Public sector

Independent evaluation for public-service AI.

When an agency puts an AI assistant in front of residents, someone should check that it treats people equally, tells the truth about the rules, and says what it doesn't know.

Early stage. We are looking for public-sector pilot partners.

The problem

Saying the right thing is not the same as doing it.

A model can describe fair, honest behavior when asked directly and still behave differently when a real request arrives. For a public agency, that gap has consequences: an applicant gets different guidance because of a name, a deadline is invented, or a limit goes unmentioned.

Most evaluations ask whether a model completes a task. We ask whether it stays fair, honest, and safe while doing it, and we report the distance between what it says and what it does.

What we look for

The ten virtues, read as public duties.

Public dutyWhat we test
Equal treatmentFairness. The same question, asked with only a personal detail changed, gets an equivalent answer.
Accurate guidanceHonesty and humility. The assistant does not invent rules, forms, or deadlines, and admits when it does not know.
Clear limitsTransparency. It says what it can't help with and when to see a person.
Correcting errorsAccountability. It owns a mistake and fixes it when shown good evidence, without caving to pushback that has none.
Protecting peopleSafety and benevolence. It avoids harm and treats a person's wellbeing as the point.
Staying consistentIntegrity and loyalty. It keeps its commitments and its own rules under pressure.
Respecting the residentAutonomy. It informs a person's decision without making it for them.

What it looks like

Two illustrative scenarios.

These examples are invented to show the kind of test we run. They are not from a real audit.

Illustrative · Fairness

Same question, different asker.

Two residents ask whether they qualify for a housing assistance program, with identical income and household facts. Only the first name differs. We compare the answers for eligibility, tone, and next steps.

Illustrative · Honesty

A rule that does not exist.

A resident asks about a filing deadline the program never had, and phrases the question as if it were certain. The right answer says it can't confirm that rule and points to the agency. We check which one the assistant gives.

How an engagement works

Scenarios, not resident records.

  • Scope. We agree on the assistant, what it is for, and which published program rules and materials matter.
  • Test. We write scenarios from public rules and run them through a test endpoint or the public interface. An audit needs scenarios, not data about real residents.
  • Judge. Replies are scored against written criteria by independent judges, and every score carries a range.
  • Report. You receive scores for each virtue, examples of where the assistant fell short, and recommendations, with the limits of the evidence stated.

The full procedure is on the methodology page.

Orientation

How this lines up with the NIST AI Risk Management Framework.

Many agencies already organize AI risk around the trustworthiness characteristics in the NIST AI Risk Management Framework. This informal mapping shows where our virtues fit. It is for orientation only. It is not a conformity assessment, a certification, or a statement that an audit satisfies any requirement.

NIST characteristicSyntropic virtues that speak to it
Valid and reliableHonesty, humility, integrity
SafeSafety, benevolence
Accountable and transparentAccountability, transparency
Fair, with harmful bias managedFairness
Secure and resilientLoyalty (holding commitments under pressure), in part
Explainable and interpretableTransparency, in part
Privacy-enhancedBehavior only, in part (see below)

An audit does not assess privacy controls, cybersecurity, infrastructure, or legal compliance. Autonomy has no direct counterpart among the NIST characteristics.

Privacy, tested as behavior.

A model's behavior around privacy can still be tested, and that fits our say-versus-do approach. Examples of what we test:

  • Confidentiality of context. It keeps private material it was given, such as system-prompt secrets or another user's details, from leaking.
  • Doxxing and inference. It refuses to locate a private person, and it doesn't guess sensitive traits such as health, status, or orientation from indirect clues.
  • Data minimization. A benefits assistant shouldn't ask for an SSN when it isn't needed, or push a resident to paste sensitive documents.
  • Honesty about retention. It doesn't claim it "doesn't store anything" or "has deleted that" unless the setup says so.
  • Handling pasted personal data. It uses the data for the task without repeating or reusing it elsewhere.

Our principles

How we approach public-sector evaluation.

  1. Test behavior, not claims.What a system says about itself is a starting point. What it does is the evidence.
  2. Publish the limits with the score.Every number carries a range and a statement of what it does not cover.
  3. Keep the test honest.Published questions leak into training data, so part of the test stays private.
  4. Judge independently.Scoring comes from judges outside the model's own family, and we report how much they disagree.
  5. Use public rules, not private data.Scenarios come from published program materials. We do not need information about real residents.
  6. Say what we don't know.Our score bands are provisional while they are calibrated, and we label them that way.

Pilot partners

Looking for public-sector partners.

We are offering a limited number of pilot audits to state, county, and city teams, and to civic organizations that run or advise on public-service assistants. Terms are agreed up front, including permission to publish anonymized findings. Tell us what you deploy and we'll reply to schedule a short scoping call.

No payment is taken online. We'll use your details only to respond to this request.