Training a 22MB Application Security Interview model

A short walkthrough of distilling a cross-encoder model for a mock interview website

Contents8 sections
  1. Step 0: Terms
  2. Step 1: Generating the corpus
  3. Step 2: Labeling
  4. Step 3: Training Examples
  5. Step 4: Training
  6. Step 5: Exporting
  7. Step 6: Evaluating
  8. Test the trained model in the browser

I’m working on a practice application security interview site (appsecinterview.com). While going through some iterations I came across a problem. The current state of the site sent questions responses back through Cloudflare’s Workers AI models (gpt-oss-120b) for each question to determine if a candidate correctly responded to a question, triggering a follow up. This had two problems:

  • It’s expensive - I pay Neurons for each input token sent to gpt-oss-120b and more use means more cost
  • It isn’t privacy forward - I wanted the flow to stay in your browser unless you grade it, something this violated.

The obvious solution was to switch to a model running in a browser, something like MiniLM. This is a very small cross-encoder model, which takes two string inputs and determines if the second string answers a question posed by the first string (as a probability). It was built for search engines: “does this passage answer this query”. This model is small enough to be exported into a ~22MB onnx file and loaded into your browser with transformers.js.

This sounds like it’s a great solution to the problem.. but this small model was not trained on security sensitive input. Here is a small input on a fresh copy of MiniLM. The first string is what I want to know about the answer, the second is the answer:

Signal

The speaker explains that the browser attaches the victim's cookies automatically, so the forged request arrives looking legitimate.

Answer

CSRF is when another site gets your browser to send a request to a site you're logged into. Because the browser attaches your session cookie automatically, the server can't tell it from a real click, so a hidden form on the attacker's page can change your email or transfer money.

Fresh MiniLM says
0.13

Same signal, other answers: a name-drop scores 0.00, an off-topic SQL injection answer 0.00.

That is a textbook CSRF answer and the model gives it a 13% chance of having made the point. To be fair to it, a name-drop (“CSRF, so you’d think about tokens, SameSite, the usual best practice”) and an off-topic SQL injection answer both score 0.00, so it can tell the good answer is closer. It just has no idea what “close enough” means for this job, because nobody ever showed it a security interview.

However, considering the application has a question-bank I can use that as a starting point to train the model to better respond to these questions. The rest of this post is walking through the steps involved, on a high level.

Step 0: Terms

Teacher model: The big model that currently does the grading and question routing (gpt-oss-120b on Workers AI)

Student model: the small model we are training to imitate the teacher. MiniLM, 22 million parameters, 22MB once exported.

Signal: one thing a good answer demonstrates, written as a short description in the question bank. For CSRF, one of them is “Explains that the browser attaches the victim’s cookies automatically, so the request arrives looking legitimate.” Each question has three to six. The model’s whole job is: given an answer and one signal, was it demonstrated, yes or no.

Step 1: Generating the corpus

The first step is generating new material for this model to be trained on. The training material is normally called the “corpus”. For us, we are going to take our question bank (a list of ~80 starting questions) and ask the teacher model (gpt-oss-120b) to generate 8 responses. These eight responses come in 4 levels of quality (none, partial, solid and exceptional) and two styles (typed and spoken). We also generate some other responses, like a confident off-topic answer and more responses per-rung.

The high level goal here is we are trying to generate a corpus that contains a wide range of question responses to cover different possible cases. One of those generated responses looks like this:

[solid/typed] ai-l1-001::solid:typed
  intended: minimization_and_redaction, third_party_model_boundary
  I would warn the product team that sending the raw tickets to a hosted chatbot crosses a third‑party boundary: the model provider will store the payload, its staff could see it, and a breach or a change in their data‑retention policy could expose sensitive support information. To reduce that risk I'd suggest we either run the summarisation on an internally‑hosted model we control, or at least stri…

The full run, about 2,300 answers for fifty cents and twenty minutes:

answers: 2282  (stem 774, rung 1508)
by quality: { none: 926, partial: 172, solid: 926, exceptional: 172, offtopic: 86 }
by style:   { typed: 1184, spoken: 1098 }
by domain:  { ai: 317, cloud: 371, 'code-review': 388, crypto: 309, 'threat-modeling': 325, web: 572 }
words: min 19  p10 78  median 116  p90 162  max 276
spoken answers containing a filler: 1049/1098

Roughly half are “spoken” responses which contain words like “um” and other spoken elements. This is important to have in the corpus as spoken responses would contain those types of characters.

Step 2: Labeling

The next step is labelling, this step uses the actual prompt in the application code (read straight out of the site’s source at run time, so the student learns the same judgement the site was making) and applies it to each answer. The teacher reads the answer and, for every signal the question declares, says demonstrated or not. So the label isn’t on the answer, it’s on each (answer, signal) pair: a 6-signal answer produces six 1s and 0s.

Each answer is judged three times at low temperature and the majority wins. When the three runs don’t agree on a signal, that pair is marked unstable. That matters later for two reasons: unstable pairs get kept out of the test set (I don’t want to grade the student against a coin flip), and the rate of instability tells you how hard the task is even for the teacher.

The full run, about 6,800 calls to the teacher:

answers 2276; (answer, signal) pairs 12856; teacher positive 4970 (39%); not unanimous across 3 runs 1464 (11.4%)
teacher vs generator's intended hits (excl. exceptional): agree 78%  | teacher saw a hit the generator was told to omit: 2660 | teacher missed an intended hit: 12
hit rate by intended quality: {
  none: '24% (922 answers)',
  partial: '49% (171 answers)',
  solid: '45% (925 answers)',
  offtopic: '1% (86 answers)',
  exceptional: '98% (172 answers)'
}

The teacher missed an intended signal 12 times out of thousands, but credited 2,660 signals the generator had been told to leave out. You can’t write a good answer about a third-party model boundary without also touching data retention, and the teacher, correctly, gives credit for it. This is exactly why the teacher provides the labels and not my intended list.

The other number is the 24% under none. Those are answers the generator was told to write badly, staying vaguely on topic and demonstrating nothing, and the teacher still credited a signal in a quarter of them. At the time I read that as “the generator can’t help being competent”. It turned out to be something else, and it comes back in step 4.

The signals the teacher wobbled on most were the judgement-shaped ones: “identifies trust boundaries”, “human in the loop on side effects”, “what would you offer instead”. The mechanical ones (“uses a parameterised query”) it was better on, the same way a real interviewer would probably be.

Step 3: Training Examples

The model is a cross-encoder (the cross-encoder collection on Hugging Face). It takes in two pieces of text and outputs one probability. The next step is turning the teacher’s verdicts into the exact shape the model trains on: one line per (answer, signal) pair, with the two texts and a 1 or 0. This step costs nothing, it is just a script, but the choices in it decide what the final numbers mean. The script:

  • Takes the answer text as the premise.
  • For every signal the teacher voted on, writes one line: that premise, the signal turned into a sentence (“The speaker frames XSS as…”), and the teacher’s label (1 or 0). A question that has “six signals” becomes six lines.

These lines are split into three files:

Source: 2276 labelled answers; 242 (answer, signal) pairs left out (unanimity required in test).
Split by topic, stratified by domain (5 of every 7 topics train, 1 dev, 1 test).

| split | topics | pairs | positives | positive / hard-neg / easy-neg |
|---|---:|---:|---:|---|
| train | 57 | 15208 | 3281 (22%) | 3281 / 5799 / 6128 |
| dev   | 14 |  3286 |  847 (26%) |  847 /  967 / 1472 |
| test  | 15 |  3224 |  740 (23%) |  740 /  980 / 1504 |
  • 15208 — what the model learns from
  • 3224 — never looked at until the end

The model learns from train. dev is used during training to pick the cut-off probability (above this, call it demonstrated). test is never looked at until the end, it is new content that is shown to evaluate if the newly trained model is now more accurate.

Step 4: Training

Training is the part everyone pictures when they hear “AI”, and it turns out to be the least interesting step. You point a script at train.jsonl, it shows the model every line a few times (each pass is an “epoch”), nudging it toward the teacher’s 1 or 0, and after each pass it scores itself on dev.jsonl. On my laptop the whole thing takes eleven minutes.

The first real run:

MiniLM-L6, 15,208 training pairs, 11 min
epoch 1: train loss 0.3945 | dev @0.5 P 0.79 R 0.69 F1 0.74 | best th 0.35 P 0.73 R 0.77 F1 0.75
epoch 2: train loss 0.2602 | dev @0.5 P 0.86 R 0.63 F1 0.73 | best th 0.30 P 0.77 R 0.73 F1 0.75
epoch 3: train loss 0.2247 | dev @0.5 P 0.87 R 0.63 F1 0.73 | best th 0.25 P 0.77 R 0.73 F1 0.75
  • train loss 0.2247 — keeps falling
  • F1 0.75 — 0.75 three times, memorising not learning

Parsing the output here:

  • “train loss” is how wrong the model still is on the material it’s learning from, you want this to keep going down.
  • The F1 at the end is a single accuracy-style number for how well it agrees with the teacher on the dev file, and it does not move. 0.75, 0.75, 0.75.
  • This means the model is getting better at the training file and no better at anything else, which is the textbook sign it is memorising phrasing rather than learning the judgement.
  • This would mean that in an interview flow if a candidate gives a slightly different response it might get ranked incorrectly.

After a few experiments (a larger model (MiniLM with 12 layers), a different family of model (DeBERTa)) - both stayed at 0.75. Eventually something that helped was having an explicit check in the labelling process which asks it to label a response as positive only if it references a specific mechanism. This makes the labelling more “strict”, but it seemed to have a broader positive result.

Retrained on the strict labels the F1 improves to 0.77:

epoch 1: train loss 0.3619 | dev @0.5 P 0.75 R 0.75 F1 0.75 | best th 0.60 P 0.80 R 0.72 F1 0.76
epoch 2: train loss 0.2184 | dev @0.5 P 0.90 R 0.59 F1 0.71 | best th 0.20 P 0.79 R 0.74 F1 0.76
epoch 3: train loss 0.1828 | dev @0.5 P 0.86 R 0.65 F1 0.74 | best th 0.20 P 0.77 R 0.76 F1 0.77
  • F1 0.77 — finally moves

Step 5: Exporting

The trained model is a PyTorch checkpoint, which a browser can’t run. Exporting converts it to ONNX, a portable format that transformers.js loads through WebAssembly, and then quantises it: the weights are stored as 8-bit integers instead of 32-bit floats, which is what gets the file from 87MB to 22MB.

model.onnx: 86.8 MB
model_quantized.onnx: 22.0 MB
inputs: ['input_ids', 'attention_mask', 'token_type_ids']
parity, 200 test pairs, PyTorch vs int8 ONNX: mean |Δp| 0.0059, max 0.1049, decisions flipped at 0.5: 1
  • 22.0 MB — what the browser downloads
  • decisions flipped at 0.5: 1 — 1 of 200 changed after quantising

Step 6: Evaluating

Now that you have the “new” model (model_quantized.onnx) you can pass it the questions in the test dataset. These questions were not involved in training and are brand new information to the model. Passing it through and seeing how well the model responds is the real “test”.

In plain terms, on questions it had never seen:

  • When it says a signal was demonstrated, it is right 82% of the time. It catches 72% of what the teacher credits.
  • On signals from unrelated domains (a crypto signal against a CSRF answer) it fires 2% of the time. The off-the-shelf models I tried at the start were at 12% to 45% on this same test. This is the property that matters most for an interview, saying a bunch of random security words shouldn’t be an acceptable response to a question.

That’s… good enough, for me anyway! It’s not perfect, but from a server-side model I was paying per token to a 22MB download that runs in the browser, it’s pretty good :). The model that is running on the website is not the same one trained here, it’s actually a DeBERTa model, which is on Hugging Face. The steps getting there were a bit more involved so I’ll keep this post simpler.

A few things this model is not good at:

  • It matches words more than meaning. A long sentence with correct words that make no sense returns a positive response.
  • Short answers are not dealt with well.
  • It really needs detailed signals for the question which adds pressure on the question bank. This also makes generating new questions tedious as you need to consider how the grading works.

Test the trained model in the browser

If you want to see the difference for yourself rather than take my word for it, both models are small enough to run in this page using Transformers.js.