Jason Stiltner
Staff AI Engineer · Agent evaluation & evaluator reliability
I build systems for knowing whether production AI actually works: the evals that catch failures, the tooling that explains them, and tests of whether the evaluators themselves can be trusted.
Staff Engineer at CharlieIQ, Greater Sum Ventures' AI lab · Shipped production AI at HCA Healthcare, the largest US hospital system · Accenture Automation CoE
Focus
- Did it fail?
- Behavioral evals that catch what unit tests cannot — When Production Disagrees and the published eval runs for the chatbot on this site.
- Why did it fail?
- Tooling that turns a failed trajectory into an explanation — the Stage 03 failure-explanation spike, whose first result did not replicate.
- Can we trust the judge?
- Falsification tests on the evaluators themselves — Evaluator Assurance, seven mechanisms on two public agent corpora, none of which survived.
Nulls, falsifications and corrections stay published — they are what makes the rest checkable.
Corpus
Evaluator Assurance: What Survived
Falsification study of seven evaluator-assurance mechanisms on two public agent corpora. None survived as a robust improvement: repetition had nothing to average over, alternate-judge pairing's headline statistic turned out to be the wrong conditional, and the one rule that passed its numeric gate cleanly — the R3 evidence-gap rule — shrank from a +11.2 pp separation to +5.8 pp (p = 0.26) once stratified by environment.
1.8% of cases had five draws that were not unanimous — 20 of 1,106 cases, 5 calls each, one judge at temperature 0, judge stage only · 0.366 / 0.465 operational overturn precision, beside 0.543 recovery pooled over both error directions · +11.2 pp → +5.8 pp the R3 evidence-gap rule's held-out separation, before and after conditioning on environment (p = 0.26)
Branch closed 2026-09-29, then corrected in place after a separate end-to-end reproduction found eleven wrong or over-strong figures — including the one this entry used to lead with. The full record, with the original wording kept beside each correction, is on the project page.
When Production Disagrees with the Architecture
A caller answered two questions when our voice agent had asked one. That small production failure exposed a larger problem: evals can preserve counterexamples without preserving what they should teach us about architecture. Production evals → architectural assumptions → prospective evidence → earned autonomy.
The evaluation machinery described in Act I is deployed. The architectural-learning loop the essay argues for is a proposal — nothing in Act III has been built.
Explaining failures
A bounded implementation of the Explain stage, run against preserved failures from this site's evaluation corpus: it generates competing explanations, requests evidence through a typed allowlist of read-only inspections, and revises its hypotheses from what it observes. The first calibration result did not survive being frozen and rerun.
4/6 → 2/30 calibration result before and after the system was frozen and rerun — 28 of 30 trials remained unresolved, with no incorrect resolution in this small sample · 0/5 held-out failures resolved, in part because the two available probes could not reach most of the evidence the system wanted to inspect
A spike, not a system: one real case moved from unresolved to a named explanation after deterministic inspection identified a grading-rule change, and one became narrower and stayed unresolved. Nothing here is deployed or in a decision path.
Chat with the Research
Grounded RAG chatbot over this site, red-team hardened. Published eval-gate numbers including the bars still unmet, and the real defects the harness found and root-caused—rate limiter, citation injection, retrieval crowding.
Grounded Commitment Learning
Multi-agent coordination through verifiable behavioral contracts. Agents commit to observable behaviors rather than inferred mental states, so a third party can check compliance without access to internal representations. The framework is formalized against Hart-Moore incomplete contracts theory; nearly all of the evidence below is simulation, and the two figures from real LLM agents are listed first.
4/4 agents better calibrated by external assessment than self-report (real LLMs) · +0.017 motivation effect in real LLMs — a null, p = 0.56, CI excludes the predicted +0.065 · 40.4% hold-up reduction in simulation (95% CI: [37.2%, 43.5%])
Corrected in place repeatedly, most recently September 2026: several published figures were traced to hardcoded chart constants rather than experiment output, and the self-selection result was falsified by an adversarial re-audit, re-measured, and then failed to reproduce in real LLM agents. The original wording and the defect are kept on the page. Only thirteen of the experiment scripts import the GCL package; everything from 19 onward is a standalone simulation.
Methods
- Empirical research. AI-accelerated hypothesis generation and experimental iteration, gated by controlled evaluation, statistical validation, and reproducible evidence. Model-generated explanations are hypotheses, not evidence.
- Mechanistic validation. Ablations and interventions that separate predictive success from causal explanation. The punishment paradox is re-derived by CI on every push; CNL's bridge ablation weekly, at full scale — both with the run behind them.
- Formal foundations. Mathematical proofs where applicable. Convergence guarantees, conservation laws, contraction mappings.
- Executable research. Research artifacts built as software: automated evaluation, reproducible experiments, CI, test coverage, inspectable results.
- Production systems. Research shaped by the constraints of systems that ship. Observability, failure recovery, deployment — GCP, Terraform, Docker, HIPAA-compliant architectures, multi-provider routing, edge inference.
Background
Path here: language, then automation, then ML.
M.A., Université Paris Cité (formerly Paris VII) — French-language graduate program
Littérature, Langues, et Civilisations des Pays Anglophones