Tag
#benchmarks
Reports tagged benchmarks.
Reports
3 reportsAug 2026
08-26 evaluation The LLM Edge in Finance Is a Division of Labor Language models can widen a financial system's field of view. That does not make them universal forecasters, portfolio managers, or sources of alpha. 08-20 agents The Skill Is Real, the Ritual Is Not Yet Measured Curated skills measurably improve coding agents where they carry procedure the model lacks; the process disciplines the frameworks sell remain unmeasured — and the nearest tests lean the other way. 08-18 agents Hard, Verifiable Agent Data: Download for Code, Build for the Rest Distillation-grade agent tasks are download-ready where checkers ship with them — code, and now terminal work; everywhere else the verifier is the scarce component, and the environment is yours to build.