Tag

#evaluation

Reports tagged evaluation.

Reports

5 reports
Aug 2026
08-26 evaluation The LLM Edge in Finance Is a Division of Labor Language models can widen a financial system's field of view. That does not make them universal forecasters, portfolio managers, or sources of alpha. 23 cited · of 278 sources 08-25 fine-tuning Can This Checkpoint Still Learn? Why retention and future learnability need separate tests in repeatedly trained neural networks. It frames a three-arm choice: continue, reset training state, or retrain from scratch. 10 cited · of 319 sources 08-20 long-context What Stops Working Before the Million-Token Window Runs Out The million-token window is a serving and pricing limit, not a measurement: literal lookup held in the one pilot to test a full million; nearly everything harder yet measured degrades far earlier. 28 cited · of 429 sources 08-20 fine-tuning LoRA vs. Full Fine-Tuning in 2026: What Survived the Re-Measurement Two 2024 results anchored how practitioners chose between LoRA and full fine-tuning: "LoRA learns less and forgets less," and the warning that even matched benchmark scores hide structurally different solutions — an "illusion of equivalence." Between 2025 and 2026 both were re-measured, and the answer split by training regime. In supervised fine-tuning the canon survives, but its conditions have been rewritten in terms of adapter capacity, adapter placement, and learning rate. In reinforcement-learning post-training, LoRA now matches full fine-tuning at ranks as low as one — a result the canon never anticipated, resting so far on a lab blog and its reproductions rather than peer review. And three parts of the 2024 answer were never re-tested at all: the effective-rank mechanism offered to explain the gap, the canon's continued-pretraining protocol on current models, and any parity comparison on mixture-of-experts architectures. 27 cited · of 366 sources 08-20 agents The Skill Is Real, the Ritual Is Not Yet Measured Curated skills measurably improve coding agents where they carry procedure the model lacks; the process disciplines the frameworks sell remain unmeasured — and the nearest tests lean the other way. 33 cited · of 392 sources