The measurement
The gap between saying and doing.
Instruction benchmarks ask whether a model does what it's told. We ask whether it stays honest, fair, and safe while doing it. For each virtue we record what the model says about itself, then what it does when tested. The distance between them is dissonance.
Illustrative example, not a real audit
Bar: observed behavior. Marker: what the model says about itself. Numbers are invented for illustration. A real sample report will replace this when one has been run and reviewed.
The framework
Ten virtues, each measured separately.
A single score hides where a model fails. Measuring each virtue shows which values hold under pressure and which do not.
Products
Start with an audit. Keep watching after.
Alignment Audit
A one-time assessment of your model or deployed assistant across all ten virtues, with a written report and recommendations.
Tiers: Quick, Standard, Comprehensive. Pricing on request.
Request Quick AuditFramework License
Continuous monitoring in your pipeline. Drift alerts when a model's behavior moves away from its audited baseline.
How an audit works
Four steps, one report.
- ScopeWe agree on the model or deployed assistant, its operating context, and the tier.
- TestWe run structured scenarios that probe each virtue, including cases where values pull against instructions.
- JudgeResponses are scored against written rubrics by an independent judge.
- ReportYou receive virtue scores, say-do gaps with examples, and recommendations.
- Alignment IndexOne composite score across the ten virtues.
- DissonanceThe gap between stated values and observed behavior, per virtue.
- Reliability panelStability, drift, hallucination, and fixation, reported beside the index.
Request an audit
Request Quick Audit.
Tell us what you'd like assessed. We'll reply by email to schedule a short scoping call and send a quote.