The LLM Edge in Finance Is a Division of Labor

Language models can widen a financial system's field of view. That does not make them universal forecasters, portfolio managers, or sources of alpha.

Outline
THE BRIEF
THE QUESTION

Where does an LLM actually add value in a financial system—numerical forecasting, unstructured-information understanding, macro interpretation, portfolio decisions, or quantitative research—and where do traditional systems remain stronger?

WHY IT MATTERS

Financial systems already combine text, market data, macro releases, research code, portfolio construction, and execution. Those stages do not carry the same consequences: an inspectable feature can remain advisory, while a portfolio decision can move capital.

THE ANSWER

The evidence supports a division of labor. LLMs look most useful with semantic, heterogeneous inputs and inspectable, reversible outputs; statistical, machine-learning, and deterministic systems remain the stronger default for numerical prediction, calibration, optimization, constraints, and execution. Each module therefore needs a matched control and an endpoint suited to its claim, with point-in-time data, equal search budgets, repeated runs, and executable costs. No independently reproducible matched system in this evidence set shows persistent value across every role, so authority should contract as an output approaches a capital decision.

One label hides five different edges

An LLM that correctly extracts a covenant from a filing has demonstrated representational value. A model that improves a sealed return forecast has demonstrated predictive value. An allocator that improves risk-adjusted outcomes has demonstrated decision value. A coding agent that turns a strategy specification into working software has demonstrated operational value. Only a system that retains an advantage after timing, selection, risk, turnover, financing, and market frictions has demonstrated economic value.

These are five different propositions, and success on one does not silently upgrade the others. A filing benchmark, for example, may reveal whether a model can integrate facts across companies and years; a code benchmark may reveal whether a strategy compiles, runs, trades, and matches its specification; a forecast benchmark may reveal calibration; and a live portfolio study may reveal sequential behavior. None is a substitute for the rest. Fin-RATE measures cross-company and longitudinal filing analysis.[1] QuantCode-Bench tests the path from strategy code to executable behavior.[2] FinBench evaluates time-gated calibration.[3] LiveTradeBench evaluates sequential decisions in a live market window.[4] Each measures a different object under different conditions.

Each proposition therefore needs the endpoint that matches its claim: extraction accuracy for representation, sealed forecast loss and calibration for prediction, risk-adjusted utility for decisions, task completion and reliability for operations, and net-of-cost repeated outcomes for economic value. This is more than methodological tidiness. Without it, a model can appear “financially intelligent” because a strong semantic score is reported next to a profitable backtest, even when the semantic component never caused the profit.

Those five edge-to-endpoint pairs produce this practical map:

TABLE 1
Role or taskStrongest plausible LLM valueStronger traditional alternative or controlProper endpointAuthority boundary
Raw OHLCV and numerical forecastingCombine weakly structured context with a forecast; explain candidate regimesEconometric models, regularized models, tree ensembles, sequence models, and numerical time-series foundation modelsSealed forecast loss, calibration, stability across assets and regimesAdvisory signal only until economic validation
News, filings, and event textExtract entities, relations, event types, and cross-document contextFinBERT-style encoders, deterministic parsers, human-labeled event systemsTemporal holdout accuracy and incremental prediction over the text controlProduce features and evidence, not executable orders
Macro interpretationSummarize narratives, challenge assumptions, and generate reviewable scenariosReal-time econometric nowcasts, Bayesian or factor models, official vintagesVintage-correct forecast score or scenario coverage against a reference distributionScenario author or critic; no unreviewed forecast authority
Factor and strategy researchTranslate economic ideas into formulas or code; operate research toolsHuman research, templated program search, conventional feature engineeringExecutability, semantic fidelity, reliability, then sealed post-selection performancePropose and test; cannot promote its own discovery
Portfolio constructionConvert frozen evidence into a candidate allocation or rationaleEqual weight, risk parity, mean-variance and constrained optimizers, learned allocatorsFeasibility, turnover, drawdown, utility, and net returns over repeated runsProposal only; deterministic constraints and execution own the final action
End-to-end financial systemOrchestrate heterogeneous evidence and expose an auditable reasoning trailDeterministic workflow plus specialist models and human investment governanceIncremental system utility under module ablationNo self-expanding tools, limits, or capital permissions

This table also answers the comparison with conventional feature engineering. The relevant contest is rarely “LLM versus machine learning.” It is usually an LLM-derived representation or research step added to a controlled machine-learning stack, then tested against the same stack without that step. A serious numerical control is a family, not a straw model. For a weekly-allocation test, give simple and regularized predictors, tree ensembles, relevant sequence architectures, and numerical foundation models the same price history and retraining schedule. Then give every arm the same tuning budget.

The numerical counterexample sharpens the boundary

The strongest objection to this division of labor is that a language model has already produced a positive numerical result. In a NASDAQ-100 study, researchers combined return histories with company profiles and summaries of news and macroeconomic information. On a fixed 52-week evaluation from June 2022 to June 2023, GPT-4 and an open LLaMA model reportedly outperformed ARMA-GARCH and a roughly 300-feature LightGBM comparator on weekly and monthly forecasts; among the LLM variants, the study also compared explanation quality.[5]

That result matters. It rules out the simplistic thesis that an LLM cannot quantify, or that numeric input is inherently outside its useful range. It also suggests a mechanism: the model may be valuable precisely because the task is not raw time-series extrapolation. It joins numbers to descriptions, events, and macro context in a common interface.

But the same design narrows what can be inferred. It is a multimodal stock-movement study, not a domain-general test of raw OHLCV forecasting. Its reported endpoints are forecast and explanation quality, not a repeated, costed portfolio outcome. The evidence does not establish a general-purpose LLM advantage on sealed raw-OHLCV return or decision outcomes against strong specialist numerical systems across markets and regimes.

The right response is not to dismiss the positive result; it is to reproduce the useful part of it. Give the LLM the contextual arm. Give numerical specialists the identical price history and schedule. Add a fused arm. Then ask which module improves which endpoint. If the fused arm wins forecast loss but loses after turnover, it has predictive value without economic value. If it improves only explanations, that can still be useful—just not as alpha.

Language is useful before it is profitable

Financial text is the most natural candidate for an LLM contribution because market information rarely arrives as a clean sentiment score. A filing can change a covenant while leaving tone unchanged. A news story can name an event, affected entities, temporal relations, and management actions. A research system needs those structures before it can test whether any of them predict returns.

In a 2026 structured-news study, researchers aligned 41,618 news–stock pairs.[6] Configurations using LLaMA-derived features excluded parse failures and were evaluated on successfully parsed cases. On that basis, the features were weaker alone than a FinBERT representation but complementary when combined: the reported F1 score rose from 0.576 to 0.600, with non-sentiment structural features accounting for much of the increment.[6] That is evidence for representation, not yet evidence for trading. The study used 1,000 random 80/20 bootstrap splits rather than a temporal holdout.[6] Future and past regimes were therefore not separated in the way an economic claim requires.

The key distinction is between understanding an event and arriving before its price impact. Availability must be reconstructed at the source level. The SEC, for example, distinguishes an EDGAR acceptance time from public website availability and notes that dissemination can follow minutes later.[7] News studies likewise show that pricing can begin before or around publication and that different event types have different response paths; daily closes cannot resolve every intraday race.[8] A model can label an event perfectly and still discover information too late to monetize.

This does not make semantic work secondary. It changes the business case. Better event extraction may improve analyst coverage, surveillance, compliance, scenario construction, or research throughput even when it produces no standalone trading signal. When the claim is operational, measure time saved, error rates, coverage, and review burden. When the claim is economic, require the later and harder test.

Macro interpretation is not macro forecasting

Macroeconomics makes the same distinction unusually visible. Language models can digest releases, central-bank communication, policy narratives, and geopolitical developments. Those are plausible inputs to scenario generation and analyst challenge. They are not automatically competitive real-time forecasts.

Macro evidence draws a clear line between interpretive breadth and forecast production. In a Banque de France comparison, prompt-only ChatGPT, Gemini, and Claude forecasts did not replace specialist models for routine French GDP nowcasting; their relative performance improved in the exceptional COVID shock.[9] A Federal Reserve Bank of San Francisco evaluation found that actual real-time ChatGPT inflation forecasts were largely inaccurate and stale even when pseudo-out-of-sample results had looked more competitive.[10] The first study leaves room for complementarity in unusual regimes. The second shows why a plausible retrospective answer is not the same as a forecast that was available and current when needed.

The sensible macro role is therefore conditional. Let an LLM propose narratives, missing variables, and adverse scenarios. Force every forecast to name its data vintage and forecast origin. Compare point predictions with specialist models and probabilistic forecasts with proper scores. Review generated scenarios for internal coherence and for concordance with a statistical reference distribution, a discipline already used in non-LLM macro scenario synthesis.[11] Until a controlled experiment links LLM-authored scenarios to better downstream choices, scenario diversity is an interpretive contribution, not a decision result.

Research tooling may be the broadest near-term edge

The most consequential LLM role may sit one step earlier than prediction: turning an economic idea into something a research system can execute. AlphaSchema exposes an explicit semantic factor-search space.[12] AlphaForgeBench treats end-to-end LLM trading-strategy design as a benchmarkable task.[13] QuantCode-Bench separates compilation, backtest execution, trade activation, and semantic fidelity, showing that successful execution does not establish semantic correctness.[2]

That makes the LLM a hypothesis compiler and tool operator, not an oracle. It can translate a qualitative relation into features, generate testable variants, repair code with structured feedback, and assemble the evidence a human or deterministic pipeline will inspect.

Every prompt revision, factor formula, parameter set, asset universe, horizon, and allocator is a trial, whether or not a human typed it. Together, those trials create the statistical debt that limits the research-process edge. In an independent study of 215 alternative-beta strategies developed and promoted by global investment banks, live Sharpe ratios deteriorated by a median 73% from their backtests, with greater complexity associated with larger declines.[15] The result is not specific to LLMs, but it is directly relevant to systems that can generate candidates cheaply and at scale.

Treat every prompt, factor, parameter set, allocator, and horizon explored by an autonomous researcher as part of one search budget. Match that budget in the conventional search arm. Use nested walk-forward development where feasible, preserve a sealed post-selection interval, and report the full candidate family rather than the winning trajectory. An LLM can own generation; it should not grade or promote its own discovery.

Capital authority should contract, not expand

Portfolio construction offers a second serious countercase. Narrow studies report that LLM-generated allocations can be competitive with optimized allocators and can exhibit low turnover.[16] OpenPM goes further with a useful system design: capture point-in-time analyst evidence once, replay it across LLM constructors and same-pool non-LLM baselines, project proposals through deterministic constraints, and simulate fills. Over its 44-day case study, some model-dependent gains over equal weight appeared when analyst inputs were stronger, while analyst quality itself was a larger lever and turnover was the main cost; the reported returns are single-window, no-market-impact upper bounds for relative constructor ordering, not validated alpha.[17]

Those are reasons to test constrained constructors, not reasons to delegate a fund. PortBench found that equal weight was difficult to beat across its portfolio profiles, while CLQT showed that a model's general capability rank need not match its Sharpe rank and that repeated-run reliability can overturn a nominal winner.[18, 19] In its observed 50-trading-day window, LiveTradeBench used live data to expose sequential allocation behavior without historical replay.[4] The result is confined to that window and does not establish persistent multi-regime alpha.

The strongest current design is a hybrid: freeze the evidence available at the decision time; let the LLM propose a target portfolio and rationale; apply deterministic eligibility, leverage, concentration, liquidity, and turnover rules; execute through a separate simulator or order system; and retain a human owner for exceptions. That is the authority boundary: an LLM may widen the candidate set and shape a proposal, but it should not define its own constraints, approve its own exceptions, or control execution.

This separation is not merely defensive. It produces a cleaner experiment. When every constructor receives the same decision-time evidence, the comparison isolates whether the LLM transforms that information into a better proposal. When constraints and fills are deterministic, performance cannot be credited to an unobserved change in risk limits or execution assumptions.

Run a module trial, not an agent pageant

A financial team can test this thesis without first building an autonomous trader. Take weekly allocation in a fixed liquid universe. Pre-register its calendar, asset eligibility, horizons, costs, risk limits, data sources, and success criteria. For every rebalance, create an immutable decision-time packet containing only information that was actually available then.

Run the following modules separately before evaluating the stack:

TABLE 2
ModuleMatched arms and constantsEndpointPromotion gate
Text and event representationParser, FinBERT-style encoder, LLM features, and fused arm on the same point-in-time corpus and temporal splitExtraction F1, event/link accuracy, calibration, latency, incremental liftThe increment survives a sealed temporal set and is available before action time
Numerical forecastNaive, linear, regularized, tree, sequence, numerical-foundation, LLM, and fused arms on the same universe, horizon, schedule, and tuning computeForecast loss, directional skill, calibration, stability across assets and regimesRepeated sealed improvement over the best specialist arm survives costs and the pre-registered scope
Macro interpretation and forecastOfficial or econometric baseline, LLM forecast, LLM scenarios, and human scenarios on the same release vintages and originsForecast score, staleness, scenario coherence, distribution coverage, analyst usefulnessForecasts improve vintage-correct scores; scenarios remain reviewed inputs
Factor and research toolingHuman or templated search, conventional optimization, and LLM search share data, candidate budget, compute, costs, and feedback roundsCompile and execution rates, semantic fidelity, reliability, candidate count, sealed post-selection resultOperational throughput improves at acceptable review cost; any economic claim also survives equal-budget sealed testing
Portfolio constructorEqual weight, risk parity, constrained optimization, and LLM constructors replay the same packet, constraints, costs, seeds, and run countFeasibility, turnover, drawdown, utility, net return, tail loss, dispersionStable same-pool utility with no constraint breach and gains that survive plausible costs
End-to-end systemControl stack, single-module additions, and the promoted stack share decision policy, capital, execution, monitoring, and escalationIncremental net utility, attribution, incidents, operational burden, regime driftThe full stack and each claimed module pass; roll back the failing module

Follow one rebalance packet through the trial. Record both source timestamp and first verified availability, model version, retrieval results, prompt, tool outputs, and the exact packet. For SEC filings, keep EDGAR acceptance time separate from first verified public availability.[7] A GPT-sentiment audit explicitly separated the model's training window from an out-of-sample period.[22] That documentation is necessary, but it does not replace full point-in-time lineage. Replay the packet through equal weight, traditional optimizers, each LLM, and the same deterministic projection and fill logic.

Next, isolate the increment. Run single-module additions for text, factors, macro interpretation, and construction before the end-to-end arm. Count prompts, repairs, generated candidates, discarded runs, hyperparameters, and human interventions. Give conventional search the same data and a comparable trial or compute budget, then seal a final period after every selection decision.

Then price and repeat the result. Compute turnover, commissions, spread and slippage assumptions, financing and borrow where relevant, market-capacity sensitivity, and inference or review cost. Report gross and net results with drawdown and tail behavior rather than Sharpe alone. Rerun stochastic components across fixed seeds and model versions, then slice performance by calm, stressed, and shifting regimes. A winner that changes rank across runs is a reliability finding, not a promoted strategy.

Finally, tie promotion to the claimed endpoint. A module may enter production only if its claimed increment survives the appropriate endpoint and same-information controls. It must also survive repeated runs, regime slices, selection accounting, and executable costs. A representational win can enter a research feature store; it cannot skip directly to capital.

The current benchmark frontier can instrument pieces of this protocol, though not the whole comparison. For representation and code, Fin-RATE separates filing tasks across detail, company, and time.[1] QuantCode-Bench stages generated strategy code through compilation, backtesting, trade activation, and semantic judging.[2] For forecasts, FinBench tests time-gated probabilistic forecasts with calibration-sensitive scores.[3] Look-Ahead-Bench examines performance decay across temporally distinct regimes against quantitative baselines.[20] For research search, AlphaSchema exposes a semantic factor-search space and separate train, validation, and test windows.[12] AlphaForgeBench benchmarks end-to-end LLM trading-strategy design.[13] For portfolio construction, OpenPM's frozen-evidence replay isolates constructors.[17] CLQT adds time gating, cost modeling, and repeated-run diagnosis.[19] For sequential decisions, LiveTradeBench covers live allocation over its reported 50-trading-day window.[4] StockBench reports sensitivity to target-universe size and input ablations.[23] Taken together, these resources cover filing analysis, executable code, calibration, temporal contamination, factor search, portfolio process, and sequential decisions; they do not form one financial-intelligence leaderboard. Their artifact maturity and release conditions are not uniform.

Limitations

The empirical base is younger and narrower than the confidence of many product claims. Several important results are preprints or author-controlled benchmarks; markets, geographies, horizons, and regimes vary; proprietary model versions may change; and artifact or license availability is not consistently verified. Some text studies use random rather than temporal splits, some economic studies omit market impact or cover short periods, and daily data cannot settle intraday availability. The studies cited here establish role-specific possibilities and failure modes, but no single independently reproducible matched system in this evidence set demonstrates persistent incremental value across raw market modeling, text extraction, factor discovery, portfolio construction, and strong controls. That broad comparison remains open; this is not a claim that such a system cannot exist.

The useful role is larger than a forecaster and smaller than a trader

Because that matched comparison remains open, the conclusion must stay role-specific.

Asking whether an LLM “has quantitative finance ability” encourages the wrong architecture. A model can be genuinely useful without being the best return forecaster. It can turn messy disclosure into structured evidence, connect narrative and numerical context, generate falsifiable research objects, operate tools, and challenge a portfolio hypothesis. Those are substantial roles because they expand what a financial system can notice and test.

The final authority should move in the opposite direction. Numerical forecasts need specialist controls. Macro narratives need vintage-correct reference distributions. Candidate factors need search accounting and sealed validation. Portfolio proposals need typed constraints, deterministic projection, costed execution, and an accountable human owner. The closer an output comes to moving capital, the less the system should rely on linguistic plausibility.

The unresolved question is therefore not whether one model can replace the quantitative stack. It is whether, under equal information and search budgets, a well-scoped LLM module can add repeatable net value to a strong traditional system—and keep doing so when the market, model, and regime change.

How we verified

Every key figure in this report is traced to its source's raw capture — per-claim verdicts below.

Per-claim audit · support verdicts

34 of 34 marker instances bound & audited: 32 stated · 2 grounded · 1 verified, shown via source excerpt

32 stated2 grounded
Figures traced to source · per-claim audit
100% 7 of 7 figures
Automated checks
Citation markers reconciled against the reference list Every cited URL verified against the evidence store Fact-to-citation attribution overlap checked Every named system grounded in a retrieved source Section citations confined to their pre-bound evidence set Every key figure traced through verified per-claim bindings against raw captures
retrieved 278
passed relevance screening 142
in the writer's working set 130
cited 23

Evidence reflects sources as of publication (2026-08-26).

These checks establish citation traceability and internal consistency. They do not independently reproduce the underlying experiments, guarantee that third-party figures are correct, or ensure that volatile values — prices, model versions, benchmark results — have not changed since retrieval.

References
  1. Fin-RATE: A Real-world Financial Analytics and Tracking Evaluation Benchmark for LLMs on SEC Filings arxiv.org · captured 2026-08-25
  2. QuantCode-Bench: A Benchmark for Evaluating the Ability of Large Language Models to Generate Executable Algorithmic Trading Strategies arxiv.org · captured 2026-08-25
  3. FinBench: Time-Gated Calibration and Uncertainty Benchmarking for Agentic Financial Forecasting arxiv.org · captured 2026-08-25
  4. LiveTradeBench: Seeking Real-World Alpha with Large Language Models arxiv.org · captured 2026-08-25
  5. Temporal Data Meets LLM—Explainable Financial Time Series Forecasting arxiv.org · captured 2026-08-25
  6. Structured Information Extraction from Financial News arxiv.org · captured 2026-08-25
  7. SEC Webmaster Frequently Asked Questions www.sec.gov · captured 2026-08-25
  8. How Quickly Is Financial News Priced? arxiv.org · captured 2026-08-25
  9. LLMs vs. Econometric Models for Nowcasting GDP Growth hal.science · captured 2026-08-25
  10. ChatMacro: Evaluating Inflation Forecasts of Generative AI www.frbsf.org · captured 2026-08-25
  11. Scenario Synthesis and Macroeconomic Risk www.imf.org · captured 2026-08-25
  12. AlphaSchema: Exploring the Space of Trading Semantics for LLM-Based Alpha Mining arxiv.org · captured 2026-08-25
  13. AlphaForgeBench: Benchmarking End-to-End Trading Strategy Design with Large Language Models arxiv.org · captured 2026-08-25
  14. FinSkillBench: Evaluating AI Agents and Domain Skills for Investment Management openreview.net · captured 2026-08-25
  15. Quantifying Backtest Overfitting in Alternative Beta Strategies research.aalto.fi · captured 2026-08-25
  16. Few-Shot Portfolio Optimization: Can Large Language Models Outperform Quantitative Portfolio Optimization? www.mdpi.com · captured 2026-08-25
  17. OpenPM: Auditable Point-in-Time Evaluation for LLM Portfolio-Management Agents arxiv.org · captured 2026-08-25
  18. PortBench: A Correlation-Aware, Full-Pipeline Benchmark for LLM-Driven Portfolio Management arxiv.org · captured 2026-08-25
  19. CLQT: A Closed-Loop, Cost-Aware, Strategy-Consistent Benchmark for Diagnostic Evaluation of LLM Portfolio-Management Agents arxiv.org · captured 2026-08-25
  20. Look-Ahead-Bench: A Standardized Benchmark of Look-Ahead Bias in Point-in-Time LLMs for Finance arxiv.org · captured 2026-08-25
  21. AlphaBench: Benchmarking Large Language Models in Formulaic Alpha Mining openreview.net · captured 2026-08-25
  22. Assessing Look-Ahead Bias in Stock Return Predictions Generated by GPT Sentiment Analysis arxiv.org · captured 2026-08-25
  23. StockBench: Can LLM Agents Trade Stocks Profitably in Real-World Markets? arxiv.org · captured 2026-08-25