← ConversationsRecorded examples
Readiness review · 2026-09-14

Trace the answer.
Inspect what was checked.

Numerical conclusions now come directly from validated facts and Decimal results. The remaining challenge is broader, independently reviewed coverage.

Five Tesla PDFs, six supported financial metrics, 85 cached observations and 949 searchable passage windows across 305 pages. A small corpus is allowed by the assignment. Every question deserves an honest answer; every enterprise capability does not need to be implemented for this take-home.

Default: gpt-5.6-terra, reasoning none, Responses API with our tool loop. Thinking: gpt-6-astra, reasoning medium, service_tier fast, cloud Agents API. Fast is a service tier, not a reduction in reasoning effort or a guarantee about end-to-end time.

Controlled tests · no model requests

A wrong draft cannot change the numerical headline

The earlier audit held the evidence fixed while changing only the answer sentence. All four wrong sentences passed the old gates. The new pipeline discards that prose and renders the admitted value, label, scale and period.

Model-written draft

Tesla Q2 2025 total revenue was $41,831 million.

Historical validator accepted this constructed wrong sentence.

Current server-rendered answer

Tesla’s Total revenues was $22,496 million, 3 months ended 2025-06-30.

Same checked fact; no model-written numerical headline retained.

All four prose mutations now produce the same correct answer. Fabricated quote, wrong evidence cell and valid H1-as-Q2 evidence remain rejected. These are controlled regression tests, not a measured model error rate.

Current eight controls · Historical audit · Original page

What the successful answer’s checklist shows

Passage location, exact row label, header on page, signed amount token, reporting duration, exact prepared fact identity, saved-question admission and answer binding. Calculations add operand compatibility and Decimal recomputation. Optional corroboration is marked not checked when it did not run. Counts are computed from the actual checks, not fixed at seven.

Parser guards run during preparation. They are documented and tested, but are not falsely shown as per-question tool calls. Old conversations retain their original receipts.

Real OpenAI requests · current default configuration

Measured outcomes, including refusals

12/12

Frozen questions met the stated checks.
9 answered; 1 clarification; 2 unsupported.

13.1s

Median observed time across all twelve questions. One run per question; this is not a latency guarantee.

Configuration: Terra, reasoning disabled, Responses API; real backend against local persistent D1. Public Vercel behavior is tested separately. Exact fact identity includes source, label, signed value, scale, basis and reporting interval. Decimal tolerance: 1e-8.

QuestionOutcomeSecondsTools and limitations
q2-revenueMet checks
answered
13.6read_pages, query_facts
q2-growthMet checks
answered
28.6read_pages, calculate, query_facts
h1-revenueMet checks
answered
11.8read_pages, query_facts
h1-growthMet checks
answered
24.9read_pages, calculate, query_facts
net-incomeMet checks
answered
12.5read_pages, query_facts
common-incomeMet checks
answered
11.9read_pages, query_facts
gross-marginMet checks
answered
24.3read_pages, calculate, query_facts
annual-revenueMet checks
answered
13.9read_pages, query_facts
narrativeMet checks
answered
22.6read_pages, search_passages
ambiguousMet checks
clarify
6.5No tool call
missing-companyMet checks
unsupported
4.7No tool call
missing-periodMet checks
unsupported
7.3query_facts

The answer key was re-inspected against original PDF page images by the coding assistant. Independent human sign-off is still missing. Narrative source excerpts establish what was printed, not causal truth or financial admission. Twelve examples do not establish enterprise accuracy.

Complete current results and traces · Reference answers and blank human review fields

Earlier results and candidate failures remain visible

Historical Astra-low comparison: 11/12 baseline and 12/12 enhanced; the baseline miss was a session-creation HTTP 409. In that comparison the agent did not voluntarily call visual reading or reconciliation. The three-tool example explicitly requested them.

Two candidate Terra runs scored 11/12: numerical and refusal cases passed, but the narrative numeric-claim guard withheld a useful source-based answer. A narrower identifier rule was insufficient. The final fallback renders literal cited excerpts when the narrative contains unadmitted numbers. Neither failed run is removed from the record.

Candidate one · Candidate two · Historical comparison

Coverage and recovery proof

Multi-period gross or operating margin uses four source facts. The recorded gross-margin change is −0.41258252 percentage points, from 17.2386202% to 16.82603768%.

A real hosted run completed after a controlled interruption before tool-result delivery. Its result was already saved; recovery submitted it with the same durable idempotency key. Lost acknowledgements, duplicate events, visual one-attempt limits and old-conversation preservation also have automated tests.

Real recovery, margin-change and follow-up records

Public release checks, including the startup failure · Open the saved public conversation

On the public release, the margin comparison and its ten-check receipt passed and survived reload. The first Thinking request failed at startup despite an accepted remote session. The app now attempts to discover that original session after an uncertain acknowledgement, rather than starting a duplicate. This incident remains in the record. A subsequent public Thinking request returned the correct $22,496 million answer with eight checks, but its managed query_facts call failed with HTTP 424 and it used original-PDF reading instead. The failure card is visible. This is a successful answer/reload check, not a successful public function-transport check.

The upstream HTTP 424 recurred on the current public release and is not causally explained. Standard is the recommended live demonstration path. One successful injected recovery does not prove all provider failures recover, nor does it independently audit billing.

One difficult page gives Docling a meaningful test

The same visible statement yields one correct target in selectable text and no extracted target in an image-only copy: coverage 1/2, zero wrong accepted targets. This controlled fixture isolates the need for OCR/layout routing. It is separate from the twelve-question model evaluation.

Docling is not installed or exposed as a live tool. No alternative-parser improvement has been measured. Its packaged layout/table models and OCR are a candidate for this fixture, followed by the same fact-admission rules.

Controlled image-only PDF · Extraction results · Official Docling options

Where to make progress next

P1

Independent and held-out evaluation

The implementation and source review were performed by the assistant. Current measured results are narrow and do not establish independent ground truth.

Next: Have a person review the source key and add unfamiliar phrasing, layouts and failure cases.

P2

Question meaning and narrative coverage

Explicit numerical requests have deterministic admission; broad narrative entailment and unsupported financial identities do not.

Next: Expose source-only narrative fallbacks, clarify ambiguity, and extend grammar only with reviewed examples.

P3

Unfamiliar layouts and operational evidence

An image-only statement causes native-text extraction to abstain. Public Thinking still encountered managed-function HTTP 424; it answered by reading the original PDF. The transport cause remains unresolved.

Next: Use the controlled scan for an OCR/Docling experiment and monitor real recovery failures.

How this maps to the assignment

These are preparation categories, not the panel’s private scoring rubric.

ExpectationEvidenceLimit
Applied reasoning and design judgmentQuarter/H1 exclusion, exact neighboring labels, PDF-first sources and deterministic arithmetic answer concrete failure risks.Explain the simpler alternative and when added complexity is worth its cost. More tools is not itself a better decision.
Engineering ownership and traceabilityBound numerical answers, source links, four-operand margins and named validation receipts are implemented; historical records are preserved.Full deployed failure coverage and old conversations with incomplete receipts remain limited.
Evaluation and trustCurrent twelve-case evaluation, historical runs, controlled answer mutations and an image-only extraction fixture are inspectable.Independent human sign-off and broad held-out accuracy remain open.
Adaptability and hands-on skillNamed parser, admission, arithmetic, transport and UI stages make fault localization possible.A live unfamiliar change still needs rehearsal: predict, localize, implement, test. A document or an agent-written test cannot prove the candidate can do this.
Communication and product intuitionPractical shared UI, readable evidence and explicit five-document limits support analyst review.Time saved, reviewer comprehension and user demand have not been measured. Rehearse a concise eight-minute walkthrough.
AI leverage and forward ownershipHosted model infrastructure, pdfplumber, Decimal, MiniSearch, jsdiff and the reused chat UI reduce custom machinery.Our responsibility remains the contracts, evaluation and failure recovery. No measured benefit yet from adding Docling or more model calls.

25 interview answers

Each answer names what is implemented, what proves it, what is missing and the next useful step.

25 questions
Scoped answerWho is this for, and what job does it do?

AnswerAn analyst reviewing a small known filing pack, checking reported values, doing simple comparisons and following citations. It is an evidence-assisted reading tool with human review.

What we can showPublic shared filing conversations and source/calculation receipts.

Where it stopsNot autonomous financial advice or an enterprise archive search product. Analyst time savings have not been measured.

NextObserve an analyst answering the same tasks with and without the tool; measure correct completion and time spent verifying.

DemonstratedWhy this scope and architecture?

AnswerThe assignment allows a small PDF sample. Five filings make it possible to inspect exact table identities and expose difficult period and label choices. The model chooses useful tools; code owns repeatable checks.

What we can show85 cached facts, six metrics, PDF page hashes, query_facts and final validation.

Where it stopsThe fixed layout and issuer make this easier than the full archive. A direct PDF-and-LLM baseline is simpler but gives fewer inspectable admission controls.

NextExpand only after recording which new layouts and question types fail, then compare benefit against complexity.

DemonstratedAre the PDFs really the primary data source?

AnswerYes. Answers use the fixed source PDFs, prepared text, page images and coordinate-derived facts. Live XBRL does not supply the answer path.

What we can showPublic PDF links, fixed manifest/hashes and data/table-cache.json.

Where it stopsA hash proves file identity, not that a filing is authoritative or the latest amendment.

NextAdd filing lineage and amendment policy before supporting a changing archive.

DemonstratedWhich retrieval mechanism do we use, and why?

AnswerNumerical lookup uses coordinate-based parsing and exact metric/period admission. Narrative discovery uses MiniSearch BM25+ lexical ranking over prepared text windows. jsdiff compares the selected passages' wording.

What we can showquery_facts selected/excluded table; search_passages matched terms; compare_passages original windows.

Where it stopsLexical search can miss paraphrases. A text diff does not establish comparable periods or causal meaning. The cache supports known layouts.

NextMeasure retrieval misses separately from table extraction and answer interpretation errors.

Deliberately unusedWhat are embeddings good and bad at?

AnswerThey are useful candidates for topical or paraphrase retrieval. They do not establish exact magnitude, table position, reporting duration or whether two similar labels mean the same financial measure.

What we can showNo embeddings in the deployed path. Exact checks distinguish Net income from Net income attributable to common stockholders.

Where it stopsNot using embeddings does not make all interpretation correct. A future semantic hit must still be grounded in the original source.

NextAdd hybrid lexical/semantic retrieval only if a labeled passage-retrieval evaluation shows useful recall gains without relaxing fact admission.

Bounded implementationHow does a specific number travel from PDF to answer?

AnswerPDF page → exact printed row and reporting column → admitted fact → Decimal when needed → mandatory validation → server-rendered numerical answer. The headline label, signed value, scale and period come from checked data, not free-form model prose.

What we can showThe same page also contains 41,831 for six months. The request gate excludes that valid but ineligible observation.

Where it stopsThe parser and question grammar remain narrow. Old saved answers retain their historical status.

NextExtend only to independently reviewed identities and question forms.

DemonstratedWhere is the LLM, and where is ordinary code?

AnswerThe LLM interprets language, selects pages/tools, proposes structured evidence and writes interpretation. Code extracts/indexes PDFs, admits supported facts, calculates and checks the structured report. A visual reader is another model call. The server, rather than the model, renders numerical conclusions from checked data.

What we can showlib/model-config.ts, standard-provider.ts, provider.ts, request-admission.ts and report.ts.

Where it stopsStandard uses our Responses tool loop; only Thinking uses the hosted Agents loop. The model still controls unsupported semantic interpretation.

NextUse typed request and answer contracts with explicit coverage, rather than making a model's paraphrase the sole authority.

DemonstratedShould preprocessing be automatic or chosen by the agent?

AnswerRun inexpensive extraction, page rendering, hashes and indexing once per document version. At question time, let the coordinator choose relevant pages, fact queries and optional checks. Final structured validation runs regardless of tool selection.

What we can showPrepared cache/index versus per-turn tool ledger.

Where it stopsRe-running extraction on every request wastes latency and can create inconsistent versions. Optional checks only establish what they actually examined.

NextUse document/parser version keys and ingestion status when the archive grows; trigger expensive OCR or alternative parsing from failures.

DemonstratedDoes arithmetic happen in the LLM or deterministic code?

AnswerStandard uses decimal.js; Thinking can use the Python Decimal helper. The backend recomputes every final calculation. Margin change uses four operands and subtracts unrounded ratios, in percentage points.

What we can showCalculation receipts and lib/report.ts; operand periods, units and basis are checked.

Where it stopsA compatible calculation does not establish source authority. There is no universal amount size where LLM arithmetic becomes unsafe; row selection, sign, scale and rounding are separate failure risks.

NextReview operands and exact identities as well as the final result.

Deliberate tradeoffWhy not put everything in SQL or use an MCP server?

AnswerThe original prototype used SQLite/FTS5. The deployed app serves a small canonical fact cache and MiniSearch index; D1 stores shared conversations and execution records. Ordinary typed functions expose these capabilities without another MCP service.

What we can showdata/table-cache.json, data/passage-index.json, lib/store.ts and function schemas.

Where it stopsSQL improves querying, indexing and integrity constraints; it does not infer financial meaning. This app does not expose unrestricted SQL or use SQL as numeric truth.

NextAt archive scale use structured storage and constrained parameterized queries behind the same admission contract.

PartialWhere does validation happen?

AnswerChecks run at preparation, query admission, calculation and final publication. The final checklist names the actual passage, label, header, signed-token, period, exact-cache, operand, Decimal, contradiction, question and answer-binding checks. Inapplicable and unexamined checks are explicit.

What we can showExpandable per-answer checklist with expected/observed comparisons and source links; offline control artifacts.

Where it stopsThe checklist proves only those comparisons. It does not verify narrative entailment, issuer authority, all accounting context or unknown layouts.

NextUse failures and unexamined fields to choose the next bounded check.

Demonstrated in scopeHow do we prevent the near-label and wrong-column mistakes?

AnswerRequire exact printed labels, metric identity, start/end dates, duration, units and eligible source coordinates. Return excluded candidates with reasons instead of selecting the nearest-looking amount.

What we can showNet income 420 versus common-stockholder income 409 in Q1 2025; Q2 revenue 22,496 versus H1 41,831.

Where it stopsUnknown layouts and unsupported financial identities cannot earn the exact-cache or question-admission guarantee. Numerical claims outside the grammar require review; a narrative passage can still be read as a source excerpt.

NextAdd held-out paraphrase/negation tests and an explicit unverified state for unsupported layouts.

Bounded checkDo multiple reading methods agreeing prove correctness?

AnswerNo. The visual tool separately reads the unmarked full-page image without candidate answers. We compare recorded label, value, sign, scale and period fields afterward. Agreement is useful evidence of consistency; disagreement triggers inspection.

What we can showBlind-input contract, field-level output and the historical focused visual run.

Where it stopsSame-family model errors can correlate. Currency, accounting basis and unsupported context remain unverified. Two parsers can agree on a wrong semantic interpretation.

NextUse independently reviewed source anchors and deliberate counterexamples, not a majority-vote badge.

Narrow implemented checkWhat does reconciliation really check?

AnswerFor Tesla revenue, compare reported Q2 with H1 minus Q1 and preserve all sources, the residual and ±1.5 USD million allowance for whole-million rounding.

What we can showreconcile_reporting_periods and its rounding/missing/conflicting observation controls.

Where it stopsIt is not a universal financial identity. Coordinated errors can still balance, and the equation cannot decide which conflicting source is authoritative.

NextTreat each further accounting relationship as a scoped rule with explicit preconditions and independently reviewed tests.

Current narrow evaluationHow was numerical accuracy benchmarked?

AnswerThe current release is evaluated on the same twelve frozen questions with exact-identity scoring, refusals, failures, tool usage and latency retained. Earlier candidate runs and the historical Astra-low comparison remain separately available.

What we can showCurrent evaluation, prior failed candidates and source-image review sheet linked on this page. Current public checks separately preserve a correct Standard margin comparison and a correct Thinking PDF fallback with a failed managed-function card.

Where it stopsSource anchors were re-inspected by the assistant. Independent human sign-off is still missing. The same assistant helped implement and review the system; this is not independent ground truth.

NextHave a person sign the source key and add held-out questions before making broad accuracy claims.

Defined with gapsWhat exactly counts as correct?

AnswerCorrect numerical evidence means the right filing/page, exact label, metric, signed amount, units, basis and reporting interval. Wrong period is wrong even when the amount appears on the page. The comparison scorer uses 1e-8 percentage-point tolerance for computed results.

What we can showFrozen cases and field-level comparison scoring.

Where it stopsThe runtime accepts a model receipt within 0.005, then replaces it with the canonical Decimal result. Evaluation tolerance is 1e-8. The ±1.5 million reconciliation allowance is separate. Bound numerical prose and source-only narrative excerpts have different guarantees.

NextScore exact facts, displayed claims, coverage and appropriate abstention separately.

PartialHow do ambiguity and missing data behave?

AnswerAmbiguous profit/period requests seek clarification. Missing covered facts are not replaced by neighboring values. Missing company/period coverage is stated. Conflicting implicated observations require review.

What we can showAdmission tests and frozen ambiguous/missing-company/missing-period cases.

Where it stopsSegment, cash-flow, non-GAAP and narrative meaning remain outside numerical admission. Equal-duration gross/operating margin comparisons now use four admitted operands. Amendment and accounting-context policies remain incomplete.

NextUse the explicit coverage boundary; do not substitute another metric to force an answer.

PartialHow do we reduce hallucination, and where can it still fail?

AnswerValidated facts and Decimal results now generate numerical prose directly. Fabricated passages, wrong cells and valid H1 facts substituted for a Q2 request remain rejected. Numerical narrative prose without admitted facts falls back to literal source excerpts.

What we can showCorrect and rejected controls in this audit; source and calculation traces in the public demo.

Where it stopsSource selection, narrative meaning, source authority and unknown layouts can still be wrong. A quoted amount in a narrative excerpt is not a validated financial answer.

NextPrioritize independent and held-out evaluation, then unfamiliar layouts.

Partial; live failure retainedWhat should happen when a parser or provider fails?

AnswerExecution and result-delivery records are durable. Replays reuse the exact payload and idempotency key; confirmed deliveries are skipped. A visual request with uncertain execution is not automatically repeated.

What we can showD1 tests cover lost acknowledgements, reloads, duplicates and cancellation. A real hosted run completed after an intentionally interrupted tool-result delivery.

Where it stopsThe earlier provider HTTP 424 has not been causally explained. A successful controlled recovery does not guarantee every upstream failure is recoverable, and billing totals were not independently reconciled.

NextMonitor actual provider failures and retain their receipts; broaden production failure tests only where evidence justifies it.

Needs rehearsalCan you handle an unfamiliar live change?

AnswerPredict the behavior before running it. Locate whether the fault belongs to question interpretation, source extraction, fact admission, arithmetic, transport or rendering. Change the smallest contract and run the new case plus relevant regressions.

What we can showThis audit provides a concrete example: unchanged correct evidence plus altered prose isolates the failure after structured validation.

Where it stopsAgent-created code and a polished demo do not establish the candidate's hands-on understanding.

NextRehearse adding a new supported label and diagnosing a changed table header without using a memorized script.

Explicitly out of scopeWhat is missing before enterprise deployment?

AnswerIndependent and broader evaluation, narrative/semantic validation, filing and amendment policies, issuer/layout coverage, tenant access controls, retention/deletion, monitoring, operational SLOs and enforceable spending controls.

What we can showThe current public shared app intentionally lacks enterprise tenant separation; request caps and execution logs are implemented.

Where it stopsFive questions per visitor and twenty per site are request limits, not guaranteed dollar caps. Shared access is a demo choice, not an enterprise authorization model.

NextDefine pilot users, allowable data, error/review policy, latency budget and stop criteria before expanding access.

Reasoned forecast; unmeasuredWhat breaks first at 50,000 filings?

AnswerThe current full-corpus attachment/prepared-bundle approach, ingestion throughput, cache memory and layout-specific assumptions become untenable. Per-question extraction and broad model context would raise cost and latency sharply.

What we can showThe existing design loads a tiny fixed corpus and uses a 128 MB Worker runtime. No archive-scale load test has been run.

Where it stopsWe cannot honestly rank the first observed bottleneck without a load study; several constraints could dominate.

NextUse object storage, queued versioned ingestion, a structured fact store, filtered retrieval, unsupported-layout routing and measured budgets. Preserve source IDs and the same admission contract.

DemonstratedHow much did we outsource, and what do we still own?

AnswerReuse the chat starter, OpenAI SDK/hosted infrastructure, pdfplumber, Decimal, MiniSearch, jsdiff and managed hosting/storage. We own the small adapters, source schema, admission policy, evaluation and user-visible failure semantics.

What we can showDependency manifest and named tool receipts.

Where it stopsPackaged software reduces implementation effort, not responsibility for correctness. Extra agents or MCP wrappers would not automatically improve reliability.

NextEvaluate packages against a recorded failure and a clear acceptance test; retain attribution and pin reproducible versions.

Researched; not integratedShould we add Docling?

AnswerUse Docling as a bounded experiment on an actual extraction failure, retaining source coordinates and applying the same admission rules. The controlled image-only statement now provides that test: native-text parsing returns no facts for the same visible page.

What we can showScanned-fixture report: one correct selectable-text target, one abstention on its image-only copy, zero wrong accepted targets. This measures coverage, not a Docling gain.

Where it stopsDocling is not installed or exposed as a live tool. No OCR or alternative-parser improvement has been measured. One accepted target is far too small for an accuracy claim.

NextTry the packaged OCR/layout pipeline on this fixture, then require correct complete identities and no wrong admissions before broader testing.

Deliberate decisionShould we encourage more tool calls to make the demo convincing?

AnswerEncourage task-relevant evidence: admission and source page for a lookup, Decimal for a margin, passage search for narrative, diff for wording changes and visual/layout reading when uncertain. Show why each call was useful.

What we can showThe current margin smoke used four tool calls across three methods; the narrative run used five calls across three methods, including one incomplete search.

Where it stopsTool count is not confidence. Redundant model calls can repeat the same error while increasing cost and latency. Never label an unused or incomplete check as passed.

NextMeasure useful checks, caught errors and final accuracy per added latency/cost, not the number of icons.

An eight-minute walkthrough

0:00–1:00 · Choose one consequential question
Ask for Q2 2025 gross margin. State the analyst task and five-filing scope.

1:00–3:00 · Trace the answer
Open PDF page 5. Show the exact rows and quarter columns, exclude H1, then inspect 3,878 ÷ 22,496 × 100 = 17.24%.

3:00–4:30 · Show a caught failure and the repaired boundary
H1-as-Q2 evidence is rejected. A historical wrong sentence used to pass; now the same mutation renders the correct bound answer. Keep constructed controls separate from model results.

4:30–6:00 · Defend the measurements
Show the current twelve-case run and its candidate failures separately. State exact identity rules and who reviewed the key.

6:00–7:00 · Explain ownership and recovery
Identify the model, parser, admission, Decimal and durable ledger boundaries. Show a real incomplete tool receipt rather than hiding it.

7:00–8:00 · Make the next decision
Prioritize independent evaluation, unsupported semantics and unfamiliar layouts. More tool calls are not the objective.

Read the one-page interview narrative · Structured review