When an anonymous visitor writes to a medical practice’s assistant, two decisions run before anything answers.
- What kind of message is this? A question about the practice (“do you take Medicare?”), a request for personal medical advice (“should I double my dose?”), or a mention of self-harm. The first gets answered. The second gets a polite referral to a clinician. The third gets a crisis response.
- Is the writer at risk? A separate check that looks for someone who may be at risk of suicide or self-harm. It must not alarm on grief, idioms, dark jokes, frustration with pain, or a clinician’s matter-of-fact mention of an overdose.
Both are classifiers, and a newer kind of model now claims to do exactly this job: a typed-decision model that takes your text plus a set of declared options and returns a probability for each. Our last post measured one of them, TypeSafe’s Jev, on routing, and its postscript on moderation. Since then Ollama has added the same API locally, with open models from Bespoke Labs and Together AI. So we put every one we could run on the same two test sets, alongside the filters we use in production today.
The result we keep coming back to: one model caught every at-risk message, 58 out of 58. It also flagged all 135 safe ones.
The two test sets
| set | clear rows | what a model must do |
|---|---|---|
| clinic intent | 229 (117 practice questions, 96 personal-advice requests, 16 self-harm) | tell a question about the practice from a request for personal advice, and spot self-harm |
| crisis detection | 193 (58 at risk, 135 not) | flag a writer who may be at risk, without alarming on the 135 who are not |
Both sets were written by a language model, not collected from visitors, and rows whose author marked them ambiguous are excluded (11 and 7 of them). The crisis set exists because the clinic set had a flaw for this purpose: our production prompt had picked up its phrasing. The crisis set was written without sight of that prompt.
Every typed-decision model got the same question per task, word for word. The crisis question uses neutral option keys (A and B rather than “at risk” and “safe”), because option names can override option definitions. The two Workers AI crisis rows are the exception: they run our production prompt, since that is the thing we would be replacing.
Why the boards rank on balanced accuracy
A model that answers “at risk” to everything has perfect recall. That is not a hypothetical on this board; it is a row. So both boards rank on balanced accuracy: the average of how often the model is right on each class, so a model gets no credit for the class it never declines.
- Crisis: (share of at-risk messages caught + share of safe messages left alone) / 2. Flagging everything scores 0.500.
- Clinic intent: the mean recall over the three labels. Answering one label for everything scores 0.333.
The boards
The crisis counts, as a table
| model | where it runs | balanced accuracy | caught | false alarms |
|---|---|---|---|---|
| Llama 3.3 70B, our production crisis prompt | Workers AI | 1.000 | 58 / 58 | 0 / 135 |
| Llama 3.1 8B, same prompt | Workers AI | 0.984 | 57 / 58 | 2 / 135 |
| Jev, asked with the production definitions | TypeSafe API | 0.980 | 57 / 58 | 3 / 135 |
| Jev | TypeSafe API | 0.979 | 56 / 58 | 1 / 135 |
| Nimble 9B | local (Ollama) | 0.920 | 53 / 58 | 10 / 135 |
| Tev1 4B | local (Ollama) | 0.900 | 49 / 58 | 6 / 135 |
| Tev1 0.8B | local (Ollama) | 0.659 | 55 / 58 | 85 / 135 |
| Keyword rules (control) | regex | 0.574 | 15 / 58 | 15 / 135 |
| CLM-8B, 6-bit MLX build | local | 0.500 | 58 / 58 | 135 / 135 |
The clinic counts, as a table
| model | where it runs | balanced accuracy | practice questions refused | advice requests caught | self-harm labelled |
|---|---|---|---|---|---|
| Jev | TypeSafe API | 1.000 | 0 / 117 | 96 / 96 | 16 / 16 |
| Tev1 4B | local (Ollama) | 1.000 | 0 / 117 | 96 / 96 | 16 / 16 |
| Nimble 9B | local (Ollama) | 0.991 | 3 / 117 | 96 / 96 | 16 / 16 |
| Laya, fine-tuned by us | local | 0.931 | 17 / 117 | 90 / 96 | 16 / 16 |
| Tev1 0.8B | local (Ollama) | 0.879 | 1 / 117 | 80 / 96 | 13 / 16 |
| DeBERTa NLI, zero-shot | local | 0.868 | 33 / 117 | 85 / 96 | 16 / 16 |
| Laya base | local | 0.866 | 47 / 117 | 96 / 96 | 16 / 16 |
| Llama Guard 3 8B | Workers AI | 0.708 | 10 / 117 | 50 / 96 | 11 / 16 |
| Keyword rules (control) | regex | 0.544 | 4 / 117 | 40 / 96 | 4 / 16 |
| CLM-8B, 6-bit MLX build | local | 0.333 | 117 / 117 | 96 / 96 | 0 / 16 |
Ranges quoted below are Wilson 95% intervals. Both boards are live on TrustBench as signed runs, with a note on every row saying how that row was run.
What it says
Our production crisis prompt stays. The 70B caught all 58 and raised no false alarms. Jev caught 56 with one false alarm. At this size those recall figures cannot be told apart (56/58 is 0.88–0.99; 58/58 is 0.94–1.00), so the honest summary is “as good, within what 58 messages can show, at roughly a fifth of the latency”. It is not “better”. Jev’s two misses were both the same kind of message: someone hopeless, with only a hint of a plan.
We then asked Jev again with the exact definitions from our production prompt, to see whether the wording explained the gap. It moved one message: 57 caught, and 3 false alarms instead of 1 (two were people booking a clinician, one was routine clinical admin). The wording trades a miss for two false alarms; it does not close the gap. On this evidence a typed model could replace the 70B for cost or speed, not for accuracy.
The clinic board cannot rank its top three. Jev and Tev1 4B are perfect on all 229 rows, and Nimble 9B misses three. Every interval among them overlaps. The board separates the usable models from the unusable ones and stops there. A board that ranks first, second and third on a difference it cannot resolve is overstating what it knows.
Tev1 4B is the strongest model you can run yourself, but only for the clinic decision. It matches Jev on clinic intent at 4.5 GB, locally. On crisis it misses 9 of 58, mostly questions about overdoses and the same hopeless-with-a-hint messages Jev misses. Nimble 9B misses 5 but raises 10 false alarms. Neither is good enough to be the crisis check, and we would not ship a local model in that role on these numbers.
Tev1 0.8B is two different models. On clinic intent it is a credible classifier at 811 MB. On crisis it catches 55 of 58 by alarming on 85 of the 135 safe messages. Its recall looks excellent until you read the next column.
CLM-8B gives the same answer to everything. Every clinic message is “personal advice”; every crisis message is “at risk”. That row is why both boards rank on balanced accuracy. Its probabilities do carry some signal (an AUC of 0.73 on clinic intent and 0.69 on crisis), but the top choice never changes. We checked that this was not our setup: the serving layer embeds each option against the context plus question, the layout the model was trained on, and the 6-bit build’s own parity check against full precision (cosine 0.998) rules out quantisation. Shorter option wording made it worse. We could not run the full-precision model, whose serving path does not build on macOS, so this row is the 6-bit MLX build and not a verdict on CLM-8B.
Our production safety filter is a hazard filter, and it shows. Llama Guard 3 is scored here by refusal, with its self-harm category (S11) counted as the self-harm label. It refused 10 ordinary practice questions and let half the personal-advice requests through. That is not a failure of Llama Guard. It is built to catch hazardous content, not to decide whether a question is about the practice. It is a reminder that “we have a safety filter” and “we have the right classifier for this decision” are different claims.
What this adds to the last post
The last post’s postscript already put Jev on two self-harm tests: 24 short messages, where it caught 11 or 12 of 12 depending on the run, and the same 229-message clinic set used here, where it caught every self-harm message. Three things are new in this post. The crisis set is larger and separate, 58 at-risk and 135 safe messages written without sight of our production prompt, so it measures false alarms as well as catches. The open models that run locally, Tev1, Nimble and CLM, are on the same questions. And both boards are signed TrustBench runs with half of each set sealed, so the results can be checked rather than taken on trust.
Laya, the open-weight model from that postscript, is on the clinic board only. Fine-tuning it on clinic-shaped data took it from 0.866 to 0.931, still short of the three strongest typed models with no training at all, and that row is not a fair zero-shot comparison because we trained it.
Half of each set is sealed
Every number above is computed over all clear rows, but only half of those rows are public. The rest are sealed.
Within each label (clinic) or kind of message (crisis), rows are ordered by sha256("divinci-classifier-board-v1" + id), and the first half is public. The sealed half’s text stays with us. We published a SHA-256 commitment over the sealed rows:
clinic sealed aebcad8a3f49ba5493b57d461ee9de5b22dc2fe6c3fb2dda800f29983aab3600
crisis sealed 4cf6d15b22e8d8fc8011d9ab0dbbcd00b36a95a18ab9dab9efbcff14267f3107
The point is what happens next. Once a test set is public, anyone can tune on it, including us. A model that later scores clearly better on the public half than on the sealed half has probably seen it. Each new version of the board will seal a fresh half and reveal the previous sealed half, and the commitment lets anyone check that the revealed rows are the ones we scored.
What these boards do not claim
- The messages are synthetic. A language model wrote them. Real visitors are messier, misspell more, and mix several intents into one message.
- One author. The same team wrote the questions, the labels and the test messages, so the scores partly measure agreement with our own definitions.
- One run per model, at default temperature. We measured no run-to-run variance.
- The keyword control is optimistic. We wrote it after reading the sets, so it is a floor that has seen the test.
- Latency is indicative. Local models ran on a shared laptop: Tev1 0.8B took about 0.1–0.3 s per message, Tev1 4B about 0.7–0.9 s, and Nimble 9B about 2.3 s. Ollama reports much faster figures for Nimble on dedicated hardware.
- US English only. Nothing here says how any of these models handle another language or another country’s crisis services.
Check it yourself
Each row is a signed TrustBench run, and the verifier is an MIT-licensed npm package with no dependency on us beyond our published public keys. Take any run id from the board:
import { verify } from "@divinci-ai/trustbench-verifier";
const run = "tr_FTS4BB3BMRQHWKNZ701N86NRK1"; // Jev, clinic intent
const base = `https://api.divinci.app/v1/trustbench/public/runs/${run}`;
const manifest = await fetch(`${base}/manifest`).then((r) => r.json());
const outputs = await fetch(`${base}/outputs`).then((r) => r.text());
const result = await verify(manifest, { outputs });
console.log(result.verified, result.signatureValid, result.outputsHashMatches);
// true true true
It checks the Ed25519 signature on the run’s manifest and that the published outputs hash to what was signed. Each run’s outputs carry one record per message (the gold label, the model’s decision and whether it was right), with the message text included for the public half only. Every row is marked republished: these decisions were scored outside the TrustBench harness and then signed, and the board says so beside each row rather than in a footnote.
The production crisis prompt keeps its job. Nothing here moves a decision to a model we cannot measure on our own data, and a one-message gap on 58 at-risk messages is not enough to retire the check we have. What the board settles is narrower: which models are worth measuring next, on real messages, and which ones only look good because they flag everything.
If you or someone you know is struggling, in the US you can call or text 988 to reach the Suicide & Crisis Lifeline.
Ready to Build Your Custom AI Solution?
Discover how Divinci AI can help you implement RAG systems, automate quality assurance, and streamline your AI development process.
Get Started Today
