# Two Papers Test Whether Language Models Can Formulate Technical Problems, Not Just Solve Them

One benchmark scores models on turning word problems into black-box optimization formulations, where formulation quality determines solution quality. The other applies Monte Carlo tree search to generating charts and analysis from tables.

- Published: 2026-08-06T05:47:01.202Z
- Canonical: https://polylog.news/ai/2026-08-06/two-papers-test-whether-language-models-can-formulate-techni
- Publisher: Polylog (AI desk)
- Section: tech
- Sources: [arXiv cs.CL](https://arxiv.org/abs/2608.02612), [arXiv cs.AI](https://arxiv.org/abs/2608.04071), [Meta AI](https://ai.meta.com/blog/genesis-mission-lawrence-berkeley-national-laboratory-segment-anything-dino/)

Two arXiv submissions this morning share a premise that separates useful engineering agents from chat assistants. The hard part of expert work is often stating the problem, not computing the answer. BBOWP-Bench evaluates large language mode…

This story is for subscribers. Read it in full at https://polylog.news/ai/2026-08-06/two-papers-test-whether-language-models-can-formulate-techni (subscription information: https://polylog.news/pricing).