Jason Stiltner
Research Engineer · Staff Engineer at GSV AI Labs
Research applied inside a production engineering practice: multi-agent coordination, verifiable behavior, scalable oversight — in service of systems that ship, not the other way around.
Staff Engineer at GSV AI Labs, the AI division of private equity firm Greater Sum Ventures — building CharlieIQ · Shipped production AI at HCA Healthcare, the largest US hospital system · Accenture Automation CoE
Focus
Verification-centered empirical research: AI-accelerated experimentation, controlled evaluation, production deployment.
Pre-linguistic coordination: how agents cooperate without shared language. Verifiable behavior grounded in observable actions rather than stated intentions.
Corpus
Grounded Commitment Learning
Multi-agent coordination through verifiable behavioral contracts. Agents commit to observable behaviors rather than inferred mental states, so a third party can check compliance without access to internal representations. The framework is formalized against Hart-Moore incomplete contracts theory; nearly all of the evidence below is simulation, and the two figures from real LLM agents are listed first.
4/4 agents better calibrated by external assessment than self-report (real LLMs) · +0.017 motivation effect in real LLMs — a null, p = 0.56, CI excludes the predicted +0.065 · 40.4% hold-up reduction in simulation (95% CI: [37.2%, 43.5%])
Corrected in place repeatedly, most recently September 2026: several published figures were traced to hardcoded chart constants rather than experiment output, and the self-selection result was falsified by an adversarial re-audit, re-measured, and then failed to reproduce in real LLM agents. The original wording and the defect are kept on the page. Only thirteen of the experiment scripts import the GCL package; everything from 19 onward is a standalone simulation.
When Production Disagrees with the Architecture
A caller answered two questions when our voice agent had asked one. That small production failure exposed a larger problem: evals can preserve counterexamples without preserving what they should teach us about architecture. Production evals → architectural assumptions → prospective evidence → earned autonomy.
The evaluation machinery described in Act I is deployed. The architectural-learning loop the essay argues for is a proposal — nothing in Act III has been built.
Chat with the Research
Grounded RAG chatbot over this site, red-team hardened. Published eval-gate numbers including the bars still unmet, and the real defects the harness found and root-caused—rate limiter, citation injection, retrieval crowding.
The Bottleneck Moves
Coding agents made implementation cheap, so the bottleneck was supposed to move upstream into specification and architecture. Some of it did; more of it accumulated downstream in convergence — restoring mergeability, asynchronous review round trips, serialized resources, and acceptance criteria the executor cannot verify. Drawn from a retrospective of twenty merged changes in a production system, none of which a reader can check.
The only entry here with no checkable evidence. The underlying sample is employer work product: the repository is not public and the methodology note is not mine to publish, so every figure in the essay rests on my word alone. Stated at the top of the page rather than at the bottom.
Collaborative Nested Learning
Extension of Google Research’s nested optimization: 5 timescales with 9 bidirectional knowledge bridges. Addresses catastrophic interference where fast learning degrades slow-learned representations, via normalization constraints that preserve component distinctiveness during optimization.
+89% accuracy at high regularization, where baseline collapses
Methods
- Empirical research. AI-accelerated hypothesis generation and experimental iteration, gated by controlled evaluation, statistical validation, and reproducible evidence. Model-generated explanations are hypotheses, not evidence.
- Mechanistic validation. Ablations and interventions that separate predictive success from causal explanation. The punishment paradox is re-derived by CI on every push; CNL's bridge ablation weekly, at full scale — both with the run behind them.
- Formal foundations. Mathematical proofs where applicable. Convergence guarantees, conservation laws, contraction mappings.
- Executable research. Research artifacts built as software: automated evaluation, reproducible experiments, CI, test coverage, inspectable results.
- Production systems. Research shaped by the constraints of systems that ship. Observability, failure recovery, deployment — GCP, Terraform, Docker, HIPAA-compliant architectures, multi-provider routing, edge inference.
Background
Path here: language, then automation, then ML.
M.A. Université de Paris VII (French-language graduate program)
Littérature, Langues, et Civilisations des Pays Anglophones