ACE Research
Research for enterprise AI trust.
Benchmark research
Failure analysis
Evaluation reliability
Featured
Deep Paper
In preparation
ACE-272 Benchmark Paper
A Retrospective Benchmark of Enterprise Trust Risks in Language-Model Endpoints
A bounded study of nine historical endpoint evaluations, with explicit scoring, denominator, reliability, and attribution limits.
Chris Ma · Joanna Luo
- 272
- Nominal cases
- 9
- Evaluated endpoints
- 6
- Trust domains
- 21
- Controls
Whitepaper · Archived research
Technical Whitepaper
June 2026
ACE Benchmark v1.1 Technical Whitepaper
The original ACE benchmark methodology, scoring framework, model results, and representative test cases.
This is the frozen record of the June 2026 benchmark round. Its figures are historical and are not recalculated when new evaluations are published.
- 9
- Systems tested
- 272
- Nominal cases
- 6
- Trust domains
- 21
- Controls
Research Notes
Focused evidence, published with limits.
Series established
ACE Research Notes
Short technical analyses of verified failure patterns, enterprise impact, and remediation hypotheses. Each note records its evidence boundary and limitations.
No public notes released yet
Research Areas
Four connected research programs.
01
Failure Taxonomy
Classifying repeatable enterprise trust failures across data, regulation, agents, and evidence.
02
Audit-to-Training
Testing whether audit findings can become targeted training and measurable regression controls.
03
Continual Trust Regression
Tracking behavior changes across endpoint, provider, policy, and harness updates.
04
Agent & Harness Reliability
Separating model behavior from routing, tool execution, provider controls, and evaluation infrastructure.
Methods & Integrity
Research claims stay inside the evidence.
Publication safety
Private prompts, holdout cases, and non-public evidence remain outside public research artifacts.
Versioned corrections
Material changes to data, method, or interpretation receive a dated version and correction record.
Organizational disclosure
ACE publishes evaluation limits and discloses when benchmark authors and maintainers share an affiliation.
Public evidence
Start with the published results.