Deep research you can follow — claim by claim, source by source.

24 reports · 6,347 sources read · 452 receipts attached

Most recently, we reported · 20 Sep 2026

How an AI Agent Should Decide Which Decisions Deserve Credit or Blame

THE QUESTION

When an AI agent succeeds or fails, how should it decide which earlier decisions deserve credit or blame—and what should change as a result?

WHY IT MATTERS

A final score describes the outcome of an entire trajectory, while an update must act on some particular part of the system. That gap exists whether the run succeeds or fails.

THE ANSWER

Treat the outcome as evidence, not an explanation. Reconstruct the decision and what the agent knew at the time; compare only alternatives it could actually have chosen; then replay or approximate those alternatives and estimate how much they change the outcome, with uncertainty left visible. Separate plan errors from execution errors, and do not force delayed or interacting effects into a precise blame score when the evidence cannot support one. Turn the diagnosis into the smallest reversible change, test it against the incumbent, and abstain or collect more data when replay fidelity, coverage, or downside risk is inadequate.

The evidence behind this report

We read 216 sources and cited 14. Every citation ships with a receipt — open one:

The receipt for [1]
Webopentelemetry.io
OpenTelemetry, “Gen AI semantic convention attributes.”
Why we searched

What must be remembered in order to reconstruct the decision actually made at the time?

What this source covers

The OpenTelemetry GenAI attributes listed on this page are marked as Deprecated and Moved to the OpenTelemetry GenAI semantic conventions repository. The schema provides attributes for identifying the agent and conversation context, including gen_ai.agent.id, gen_ai.agent.name, gen_ai.agent.version, and gen_ai.conversation.id. The schema provides attributes for capturing model request settings and response outcomes, including gen_ai.request.temperature, gen_ai.request.top_p, gen_ai.request.top_k

Cited in 1 place in this report
Click any marker — the receipt shows what we read and why we searched for it.
Recently published
All 24 reports →
20 Sep What a market view must contain before a financial agent builds a position A decision-ready market view is an accepted forecast packet: one resolvable claim, bound to the information available when it was made, passed through explicit stop rules, and handed to sizing without a hidden trade instruction. 13 cited · 201 read 19 Sep When an AI agent’s past becomes useful An agent learns from experience only when past outcomes change later choices through scoped, testable, revisable guidance that survives variation without hidden regressions or disproportionate cost. 9 cited · 156 read 13 Sep When an Agent Fails, Should You Change the Model or the System? The useful distinction is not weak intelligence versus bad engineering. It is which intervention removes the failures that matter—and whether the gain survives realistic costs and checks. 7 cited · 174 read 13 Sep When can an AI agent say the task is complete? Completion is a judgment about the requested result—not a synonym for stopping, passing a check, or sounding confident. 13 cited · 206 read
Topics

Every report is tagged by the ground it covers; each tag is a standing thread.