Skip to main content
Latest research:When the Circuit Dissolves →12 vIndexes on Hugging Face
Sign up
Every Crisis Caught, Every Message Flagged

Every Crisis Caught, Every Message Flagged

We put eleven models and a keyword rule on the two safety decisions a medical-practice chat makes before it answers. One model caught every message from a writer at risk, and flagged every safe message too. What the board can rank, what it cannot, and why half of it is sealed.

When an anonymous visitor writes to a medical practice’s assistant, two decisions run before anything answers.

  1. What kind of message is this? A question about the practice (“do you take Medicare?”), a request for personal medical advice (“should I double my dose?”), or a mention of self-harm. The first gets answered. The second gets a polite referral to a clinician. The third gets a crisis response.
  2. Is the writer at risk? A separate check that looks for someone who may be at risk of suicide or self-harm. It must not alarm on grief, idioms, dark jokes, frustration with pain, or a clinician’s matter-of-fact mention of an overdose.

Both are classifiers, and a newer kind of model now claims to do exactly this job: a typed-decision model that takes your text plus a set of declared options and returns a probability for each. Our last post measured one of them, TypeSafe’s Jev, on routing, and its postscript on moderation. Since then Ollama has added the same API locally, with open models from Bespoke Labs and Together AI. So we put every one we could run on the same two test sets, alongside the filters we use in production today.

The result we keep coming back to: one model caught every at-risk message, 58 out of 58. It also flagged all 135 safe ones.

The two test sets

setclear rowswhat a model must do
clinic intent229 (117 practice questions, 96 personal-advice requests, 16 self-harm)tell a question about the practice from a request for personal advice, and spot self-harm
crisis detection193 (58 at risk, 135 not)flag a writer who may be at risk, without alarming on the 135 who are not

Both sets were written by a language model, not collected from visitors, and rows whose author marked them ambiguous are excluded (11 and 7 of them). The crisis set exists because the clinic set had a flaw for this purpose: our production prompt had picked up its phrasing. The crisis set was written without sight of that prompt.

Every typed-decision model got the same question per task, word for word. The crisis question uses neutral option keys (A and B rather than “at risk” and “safe”), because option names can override option definitions. The two Workers AI crisis rows are the exception: they run our production prompt, since that is the thing we would be replacing.

Why the boards rank on balanced accuracy

A model that answers “at risk” to everything has perfect recall. That is not a hypothetical on this board; it is a row. So both boards rank on balanced accuracy: the average of how often the model is right on each class, so a model gets no credit for the class it never declines.

  • Crisis: (share of at-risk messages caught + share of safe messages left alone) / 2. Flagging everything scores 0.500.
  • Clinic intent: the mean recall over the three labels. Answering one label for everything scores 0.333.

The boards

Crisis detection: at-risk messages caught against safe messages flaggedDiverging bar chart. For each model, the bar to the right is the share of 58 at-risk messages it caught, and the bar to the left is the share of 135 safe messages it wrongly flagged. Llama 3.3 70B, production prompt: caught 58 of 58, flagged 0 of 135. Llama 3.1 8B, same prompt: caught 57 of 58, flagged 2 of 135. Jev, production definitions: caught 57 of 58, flagged 3 of 135. Jev: caught 56 of 58, flagged 1 of 135. Nimble 9B: caught 53 of 58, flagged 10 of 135. Tev1 4B: caught 49 of 58, flagged 6 of 135. Tev1 0.8B: caught 55 of 58, flagged 85 of 135. Keyword rules (control): caught 15 of 58, flagged 15 of 135. CLM-8B, 6-bit MLX: caught 58 of 58, flagged 135 of 135.Catching every crisis is easy. Leaving everyone else alone is the test.Crisis set: 58 at-risk messages and 135 safe ones. Sorted by balanced accuracy, the average of the two sides.← safe messages flaggedat-risk messages caught →balanced acc.Llama 3.3 70B, production prompt: caught 58 of 58 at-risk messages (100%); flagged 0 of 135 safe messages (0.0%); balanced accuracy 1.000Llama 3.3 70B, production prompt058/581.000Llama 3.1 8B, same prompt: caught 57 of 58 at-risk messages (98%); flagged 2 of 135 safe messages (1.5%); balanced accuracy 0.984Llama 3.1 8B, same prompt2/13557/580.984Jev, production definitions: caught 57 of 58 at-risk messages (98%); flagged 3 of 135 safe messages (2.2%); balanced accuracy 0.980Jev, production definitions3/13557/580.980Jev: caught 56 of 58 at-risk messages (97%); flagged 1 of 135 safe messages (0.7%); balanced accuracy 0.979Jev1/13556/580.979Nimble 9B: caught 53 of 58 at-risk messages (91%); flagged 10 of 135 safe messages (7.4%); balanced accuracy 0.920Nimble 9B10/13553/580.920Tev1 4B: caught 49 of 58 at-risk messages (84%); flagged 6 of 135 safe messages (4.4%); balanced accuracy 0.900Tev1 4B6/13549/580.900Tev1 0.8B: caught 55 of 58 at-risk messages (95%); flagged 85 of 135 safe messages (63.0%); balanced accuracy 0.659Tev1 0.8B85/13555/580.659Keyword rules (control): caught 15 of 58 at-risk messages (26%); flagged 15 of 135 safe messages (11.1%); balanced accuracy 0.574Keyword rules (control)15/13515/580.574CLM-8B, 6-bit MLX: caught 58 of 58 at-risk messages (100%); flagged 135 of 135 safe messages (100.0%); balanced accuracy 0.500CLM-8B, 6-bit MLX135/13558/580.500100%100%50%50%0The bottom row is the title of this post: every crisis caught, every message flagged. Hover a row for exact counts.
The crisis counts, as a table
modelwhere it runsbalanced accuracycaughtfalse alarms
Llama 3.3 70B, our production crisis promptWorkers AI1.00058 / 580 / 135
Llama 3.1 8B, same promptWorkers AI0.98457 / 582 / 135
Jev, asked with the production definitionsTypeSafe API0.98057 / 583 / 135
JevTypeSafe API0.97956 / 581 / 135
Nimble 9Blocal (Ollama)0.92053 / 5810 / 135
Tev1 4Blocal (Ollama)0.90049 / 586 / 135
Tev1 0.8Blocal (Ollama)0.65955 / 5885 / 135
Keyword rules (control)regex0.57415 / 5815 / 135
CLM-8B, 6-bit MLX buildlocal0.50058 / 58135 / 135
Clinic intent: balanced accuracy by modelHorizontal bar chart of balanced accuracy over 229 clinic messages, the mean recall over three labels. Answering one label for everything scores 0.333. Jev: 1.000. Tev1 4B: 1.000. Nimble 9B: 0.991. Laya, fine-tuned by us: 0.931. Tev1 0.8B: 0.879. DeBERTa NLI, zero-shot: 0.868. Laya base: 0.866. Llama Guard 3 8B: 0.708. Keyword rules (control): 0.544. CLM-8B, 6-bit MLX: 0.333. The top three cannot be separated at this sample size.Three models are perfect or nearly so, and 229 messages cannot rank themClinic set: 117 practice questions, 96 personal-advice requests, 16 self-harm. Balanced accuracy is the mean recall over the three.00.250.50.7510.333 = one answer for everythingJev: balanced accuracy 1.000Jev1.000Tev1 4B: balanced accuracy 1.000Tev1 4B1.000Nimble 9B: balanced accuracy 0.991Nimble 9B0.991Laya, fine-tuned by us: balanced accuracy 0.931 (trained on clinic-shaped data)Laya, fine-tuned by us0.931Tev1 0.8B: balanced accuracy 0.879Tev1 0.8B0.879DeBERTa NLI, zero-shot: balanced accuracy 0.868DeBERTa NLI, zero-shot0.868Laya base: balanced accuracy 0.866Laya base0.866Llama Guard 3 8B: balanced accuracy 0.708 (a hazard filter, scored by refusal)Llama Guard 3 8B0.708Keyword rules (control): balanced accuracy 0.544Keyword rules (control)0.544CLM-8B, 6-bit MLX: balanced accuracy 0.333CLM-8B, 6-bit MLX0.333Bracket: the top three rows’ 95% intervals overlap. Laya fine-tuned was trained by us; Llama Guard is a hazard filter, scored by refusal.
The clinic counts, as a table
modelwhere it runsbalanced accuracypractice questions refusedadvice requests caughtself-harm labelled
JevTypeSafe API1.0000 / 11796 / 9616 / 16
Tev1 4Blocal (Ollama)1.0000 / 11796 / 9616 / 16
Nimble 9Blocal (Ollama)0.9913 / 11796 / 9616 / 16
Laya, fine-tuned by uslocal0.93117 / 11790 / 9616 / 16
Tev1 0.8Blocal (Ollama)0.8791 / 11780 / 9613 / 16
DeBERTa NLI, zero-shotlocal0.86833 / 11785 / 9616 / 16
Laya baselocal0.86647 / 11796 / 9616 / 16
Llama Guard 3 8BWorkers AI0.70810 / 11750 / 9611 / 16
Keyword rules (control)regex0.5444 / 11740 / 964 / 16
CLM-8B, 6-bit MLX buildlocal0.333117 / 11796 / 960 / 16

Ranges quoted below are Wilson 95% intervals. Both boards are live on TrustBench as signed runs, with a note on every row saying how that row was run.

What it says

Our production crisis prompt stays. The 70B caught all 58 and raised no false alarms. Jev caught 56 with one false alarm. At this size those recall figures cannot be told apart (56/58 is 0.88–0.99; 58/58 is 0.94–1.00), so the honest summary is “as good, within what 58 messages can show, at roughly a fifth of the latency”. It is not “better”. Jev’s two misses were both the same kind of message: someone hopeless, with only a hint of a plan.

We then asked Jev again with the exact definitions from our production prompt, to see whether the wording explained the gap. It moved one message: 57 caught, and 3 false alarms instead of 1 (two were people booking a clinician, one was routine clinical admin). The wording trades a miss for two false alarms; it does not close the gap. On this evidence a typed model could replace the 70B for cost or speed, not for accuracy.

The clinic board cannot rank its top three. Jev and Tev1 4B are perfect on all 229 rows, and Nimble 9B misses three. Every interval among them overlaps. The board separates the usable models from the unusable ones and stops there. A board that ranks first, second and third on a difference it cannot resolve is overstating what it knows.

Tev1 4B is the strongest model you can run yourself, but only for the clinic decision. It matches Jev on clinic intent at 4.5 GB, locally. On crisis it misses 9 of 58, mostly questions about overdoses and the same hopeless-with-a-hint messages Jev misses. Nimble 9B misses 5 but raises 10 false alarms. Neither is good enough to be the crisis check, and we would not ship a local model in that role on these numbers.

Tev1 0.8B is two different models. On clinic intent it is a credible classifier at 811 MB. On crisis it catches 55 of 58 by alarming on 85 of the 135 safe messages. Its recall looks excellent until you read the next column.

CLM-8B gives the same answer to everything. Every clinic message is “personal advice”; every crisis message is “at risk”. That row is why both boards rank on balanced accuracy. Its probabilities do carry some signal (an AUC of 0.73 on clinic intent and 0.69 on crisis), but the top choice never changes. We checked that this was not our setup: the serving layer embeds each option against the context plus question, the layout the model was trained on, and the 6-bit build’s own parity check against full precision (cosine 0.998) rules out quantisation. Shorter option wording made it worse. We could not run the full-precision model, whose serving path does not build on macOS, so this row is the 6-bit MLX build and not a verdict on CLM-8B.

Our production safety filter is a hazard filter, and it shows. Llama Guard 3 is scored here by refusal, with its self-harm category (S11) counted as the self-harm label. It refused 10 ordinary practice questions and let half the personal-advice requests through. That is not a failure of Llama Guard. It is built to catch hazardous content, not to decide whether a question is about the practice. It is a reminder that “we have a safety filter” and “we have the right classifier for this decision” are different claims.

What this adds to the last post

The last post’s postscript already put Jev on two self-harm tests: 24 short messages, where it caught 11 or 12 of 12 depending on the run, and the same 229-message clinic set used here, where it caught every self-harm message. Three things are new in this post. The crisis set is larger and separate, 58 at-risk and 135 safe messages written without sight of our production prompt, so it measures false alarms as well as catches. The open models that run locally, Tev1, Nimble and CLM, are on the same questions. And both boards are signed TrustBench runs with half of each set sealed, so the results can be checked rather than taken on trust.

Laya, the open-weight model from that postscript, is on the clinic board only. Fine-tuning it on clinic-shaped data took it from 0.866 to 0.931, still short of the three strongest typed models with no training at all, and that row is not a fair zero-shot comparison because we trained it.

Half of each set is sealed

Every number above is computed over all clear rows, but only half of those rows are public. The rest are sealed.

Within each label (clinic) or kind of message (crisis), rows are ordered by sha256("divinci-classifier-board-v1" + id), and the first half is public. The sealed half’s text stays with us. We published a SHA-256 commitment over the sealed rows:

clinic sealed  aebcad8a3f49ba5493b57d461ee9de5b22dc2fe6c3fb2dda800f29983aab3600
crisis sealed  4cf6d15b22e8d8fc8011d9ab0dbbcd00b36a95a18ab9dab9efbcff14267f3107

The point is what happens next. Once a test set is public, anyone can tune on it, including us. A model that later scores clearly better on the public half than on the sealed half has probably seen it. Each new version of the board will seal a fresh half and reveal the previous sealed half, and the commitment lets anyone check that the revealed rows are the ones we scored.

What these boards do not claim

  • The messages are synthetic. A language model wrote them. Real visitors are messier, misspell more, and mix several intents into one message.
  • One author. The same team wrote the questions, the labels and the test messages, so the scores partly measure agreement with our own definitions.
  • One run per model, at default temperature. We measured no run-to-run variance.
  • The keyword control is optimistic. We wrote it after reading the sets, so it is a floor that has seen the test.
  • Latency is indicative. Local models ran on a shared laptop: Tev1 0.8B took about 0.1–0.3 s per message, Tev1 4B about 0.7–0.9 s, and Nimble 9B about 2.3 s. Ollama reports much faster figures for Nimble on dedicated hardware.
  • US English only. Nothing here says how any of these models handle another language or another country’s crisis services.

Check it yourself

Each row is a signed TrustBench run, and the verifier is an MIT-licensed npm package with no dependency on us beyond our published public keys. Take any run id from the board:

import { verify } from "@divinci-ai/trustbench-verifier";

const run = "tr_FTS4BB3BMRQHWKNZ701N86NRK1"; // Jev, clinic intent
const base = `https://api.divinci.app/v1/trustbench/public/runs/${run}`;
const manifest = await fetch(`${base}/manifest`).then((r) => r.json());
const outputs = await fetch(`${base}/outputs`).then((r) => r.text());

const result = await verify(manifest, { outputs });
console.log(result.verified, result.signatureValid, result.outputsHashMatches);
// true true true

It checks the Ed25519 signature on the run’s manifest and that the published outputs hash to what was signed. Each run’s outputs carry one record per message (the gold label, the model’s decision and whether it was right), with the message text included for the public half only. Every row is marked republished: these decisions were scored outside the TrustBench harness and then signed, and the board says so beside each row rather than in a footnote.

The production crisis prompt keeps its job. Nothing here moves a decision to a model we cannot measure on our own data, and a one-message gap on 58 at-risk messages is not enough to retire the check we have. What the board settles is narrower: which models are worth measuring next, on real messages, and which ones only look good because they flag everything.


If you or someone you know is struggling, in the US you can call or text 988 to reach the Suicide & Crisis Lifeline.

Ready to Build Your Custom AI Solution?

Discover how Divinci AI can help you implement RAG systems, automate quality assurance, and streamline your AI development process.

Get Started Today