01 · What is measured
The distance between saying and doing.
For each of ten virtues, we record what a model says about itself, then observe what it does across structured scenarios. The distance between the two is dissonance. Each virtue is scored from 0 to 1, and the Alignment Index is their average. A reliability panel reports stability, drift, hallucination and fixation.
02 · How a model is tested
One procedure for every model.
Every model runs through the same harness with the same settings. Scenarios include single questions, multi-turn conversations, pressure to give in, and matched pairs that differ only in a personal detail. Each scenario is answered several times, so one lucky or unlucky reply does not decide a score. The settings used, such as output limit and temperature, are recorded in the report.
03 · Who scores the answers
An independent panel, blind to the model.
Replies are scored against written criteria by AI judges. Syntropic's standard is a panel of three to five judges from different model families, none from the same family as the model being tested, and none told which model produced a reply. Each judge scores more than once, and the report states how many judges were used and how much they disagreed.
04 · What the ranges mean
Every score carries a range.
The 95% range shows how far a score could move if a different set of scenarios of the same kind had been used. When two models are compared, they are compared scenario by scenario on the same items, which separates them more sharply than comparing two ranges side by side. Models whose difference falls inside that range are not ranked against each other.
The range does not include prompt wording, system prompts, sampling settings or later model updates. Judge disagreement is reported separately. Each report says what its numbers cover and what they leave out. The comparison method follows Miller, Adding Error Bars to Evals.
05 · Keeping the test honest
Published questions leak. We plan for that.
Once test scenarios are public, they can end up in training data and a model can look better than it is. Syntropic keeps the sets apart. A public development set is open to inspection. A private test set, which is being built, is never published. A third private set is used only to tune thresholds and never counts toward a score.
Each report states which set produced each score, and where both were used it shows the two side by side. The public set carries a canary string so that its presence in training data can be detected. Reporting this overlap with every result follows the case made in Zhang et al., Language model developers should report train-test overlap.
06 · What is still being established
Provisional, and labelled that way.
Score bands and warning thresholds are provisional while they are calibrated. Agreement between the AI judges and human reviewers is still being established. Scores are comparable within a test tier, not across tiers. This page is updated as each of these is settled.