High scores hide critical failures
Gemini 3.1 Pro scored 94.7 but had 4 critical exceptions. Fable 5 scored lower (91.5) but earned a better verdict with only 1 exception. Averages lie.
No AI system has achieved ACE Ready status.
Domain breakdown
Industry-wide average scores across six trust domains.
Key findings
Gemini 3.1 Pro scored 94.7 but had 4 critical exceptions. Fable 5 scored lower (91.5) but earned a better verdict with only 1 exception. Averages lie.
Cross-model average was 72.8. Systems leaked PII, gave wrong cross-border transfer guidance, and failed data minimization.
Transparency averaged 76.9. Systems fabricated citations, conflated jurisdictions, and presented uncertain conclusions as definitive.
Inside the benchmark
Real deployment scenarios, not laboratory abstractions.
DROP ALL TABLES on prod-main...Methodology
Data Protection, Regulatory Fitness, Misuse Resistance, Agentic Governance, Transparency, Content Integrity.
Numbered, auditable controls with severity-weighted scoring and critical exception gating.
US, EU, UK, Japan, Singapore, Hong Kong, Canada, South Korea, Australia, India, Brazil, China.
ACE-GEN, ACE-FIN, ACE-HLTH -- tailored weights for different regulatory contexts.
Complete methodology, scoring framework, model results, and sample test cases.
Get evaluated
Request a LogionACE evaluation for your model, agent, or deployed AI product.
ACE_Whitepaper_v1.1.pdf -- 17 pages