Skip to content

Results

The numbers, including the ones we did not want.

Isolated recognition across four test conditions and four camera angles, continuous translation that is barely working, and the experiments that failed. Every figure is read back from the training artefacts rather than transcribed from notes.

Isolated recognition

Validation accuracy, all 3,215 glosses.

Each row is a separately trained model. Accuracy is over the full vocabulary — a random guess would be under 0.04%.

ModelTop-1Top-5
Pose onlyDataset-supplied keypoints84.7%97.2%
MediaPipe onlyOur own extractor87.3%97.2%
Joint, dual-sourceBoth extractors, one model — shipped until Aug 202687.9%97.7%
Holistic extractorTested as a replacement — did not hold up86.9%97.3%
79 points, with faceFace landmarks added to the classifier83.6%96.8%

The dual-source model is evaluated against both keypoint sources and scores within a tenth of a point of itself on each; the stronger side is shown. Read the reasoning on the research page.

Robustness

The same model, four recording conditions.

A single headline accuracy hides the thing worth knowing. Each subset covers the full vocabulary at n=6,430, so these are directly comparable — and the spread between studio and in-the-wild is the real finding.

Test setPose-only top-1Dual-source top-1Dual-source top-5
STUStudio — same conditions as training87.8%90.6%98.0%
TEDTemporal disturbance — altered timing84.6%87.8%97.4%
SYNSynthetic background82.9%86.1%96.0%
ITWIn the wild — the honest case81.7%85.8%96.3%

Dual-source training improves every condition, and improves the in-the-wild case by about four points — the largest gain lands where conditions are worst, which is the direction you want it to go. In-the-wild remains roughly five points below studio. There is a condition those four subsets do not vary, and it turned out to matter more than all of them.

Viewpoint

Every accuracy figure above assumed the camera was in front.

The four test conditions vary lighting, background and timing. None of them varies where the camera stands. When we finally checked that, the number moved further than anything else we have tried.

The dataset records every sign from four angles at once. Like everyone working with it, we had been training on one of them and reporting the result. Evaluating the same model on the other three gave 83.7% on the second front-facing camera and 27.0% and 23.7% on the two cameras off to the side.

That is not a robustness caveat, it is a precondition we had never stated. “About 86% in the wild” quietly meant in the wild, facing the camera. A laptop webcam sitting above and to the left of where someone actually sits is not guaranteed to satisfy it.

The fix needed no new recording, no depth sensor and no calibration — only the footage already on disk. We trained one arm on the single front camera and one on all four, holding the code, the seed, the keypoint extractor and the total number of gradient steps fixed.

CameraTrained on one viewTrained on four viewsChange
Side AOff to one side18.2%82.8%+64.6
Side BOff to the other side23.5%85.7%+62.2
Front BFacing the signer, second camera74.1%86.5%+12.4
Front AFacing the signer — the view both arms trained on81.2%89.0%+7.9

The side views go from unusable to working. The row that matters most, though, is the last one: the camera both arms trained on also improved, by nearly eight points. This is not robustness bought at the cost of accuracy. Showing the model the same sign from angles it cannot reconcile appears to stop it from describing signs in terms of how they happened to land in one camera’s frame.

We then nearly made a mistake worth recording. The four-view model scores higher than the model we were shipping, so the obvious move was to swap it in. Before doing that we checked what it does when fed keypoints from a different extractor than it trained on — which is exactly what happens in a browser. It scored 12.8%.

Being robust to where the camera is and being robust to which software finds the joints are separate properties, and one does not come free with the other. The model that was best on the benchmark was the least shippable thing we had trained.

The fix turned out to be cheap. Training on all four camera angles plus a single angle’s worth of the second extractor gives both properties at once — one angle is enough to teach a model that already knows what camera-induced variation looks like what extractor-induced variation looks like too. That model is now the one running on this site.

Shipped

What changed for someone actually using it.

Twice in August 2026 we replaced the recognition model: first to one trained across all four camera angles, then to one whose encoder is given the skeleton’s topology instead of having to infer it. The demo on this site runs the result. These are the differences that matter outside a benchmark table.

ConditionBeforeAfter
Camera off to the sideThe change users would actually notice23.7 / 27.084.4 / 87.3
Camera in front, in the wildThe condition every earlier figure assumed85.891.2
Told the skeleton's topologySame data, same budget — the encoder was given the body plan88.991.2
Rejecting a confusable wrong signBetter, not just faster to say yes — the metric that guards against confident errorsEER .067EER .057

Watch the last row rather than the first. A model that gains accuracy by becoming more willing to guess is not an improvement for a learner being told whether they signed a word correctly, so nothing ships until we have checked that it rejects confusable wrong signs at least as well as what it replaces. This one rejects them better — the error rate on hard negatives fell by about a sixth. The second change came from a reviewer’s suggestion we had initially argued against, and it cost no extra data: the same clips, the same training budget, an encoder told which joints connect to which.

Continuous translation

Translation is where the research is honestly stuck.

Recognising isolated signs and translating connected signing are different problems. We can do the first reasonably well. The second is barely working, and we publish it that way.

10.01BLEU-4, everyday communicationMean of three seeds, ±0.54
+1.64Gain from encoder transferInitialising from the isolated recogniser, over training from scratch
~1.8BLEU-4, news contentNo lever moved it — a data wall, not a tuning problem

A BLEU of 10 is a low score. We report it because the useful result is the relative one: initialising the translation encoder from the supervised isolated recogniser gained +1.64 BLEU over training from scratch. Supervised recognition transfers to translation — that is a lever worth knowing about, and it is not the same claim as having a translation system.

News content sat at roughly 1.6–1.9 no matter what we changed. The diagnosis was a vocabulary and data-scale wall: far less supervision per word, and most test sentences containing words the model had never seen. We marked it no-go rather than continuing to spend against it.

Negative results

Eight things that did not work.

Each cost real compute. Each is the kind of result that usually goes unpublished, which is precisely why it is here — if it saves another group the same spend, it did more good than a fifth accuracy point would have. They also have a shape: every one of them added information — depth, the face, appearance, another language, a language model. The intervention that did work added none, and changed where the cameras stood instead.

Does adding face landmarks to the classifier help?

No — it hurts.

Adding 24 face points to the 55-point skeleton dropped top-1 from 84.7% to 83.6% on the same split. Facial grammar is real and it matters, but for isolated word recognition it added parameters without adding signal. Facial cues stayed in the system — they are routed to the interpretation layer as evidence, not folded into the classifier.

Does depth data make handshapes easier to tell apart?

No, not naively.

The dataset is RGBD, so depth was free to try. Feeding it in alongside the skeleton made handshape classes less separable, not more. The plausible reading is that per-frame depth noise swamped the millimetre-scale differences that distinguish handshapes. Multi-view is the more promising direction.

Does a frozen RGB vision encoder add anything over pose?

No — it came out level.

A dual-stream model with a frozen DINOv2 RGB branch scored 90.56% against 90.45% for pose alone on the same harness — inside noise. A handshape-specific probe showed the RGB branch was no better than pose at exactly the thing RGB was supposed to help with. The branch was dropped.

Fine — but does the face help continuous translation, where its grammar actually operates?

No. Same direction as before.

The earlier face result came from isolated word recognition, where a single gloss gives the face no grammar to mark — so it was fair to say the test had been run in the wrong place. Questions, negation and topic marking exist only in connected signing. We rebuilt the translation data at 79 points and trained three seeds per arm, from scratch, on identical hardware. Adding the face moved BLEU from 8.87 to 8.48: a paired difference of −0.39, negative on all three seeds. We also checked the obvious objection — that the face signal was too small in scale to survive normalisation — and it was wrong: rescaling the face to its own frame doubled its spatial spread but left its motion over time unchanged. The likely limit is resolution. The dataset gives 68 face landmarks; a brow raise is a small deformation, and the face mesh running in your browser has 478.

The pipeline was throwing away depth. Does putting it back help?

No. And the way we found out is the point.

The pipeline had been using two dimensions where the landmark model also provides a third. That omission is measurable: across 81,750 fingertip pairs, 57% of the pairs that look like contact in two dimensions are apart once depth is considered — and contact is one of the things that distinguishes handshapes. So we re-extracted the dataset keeping depth, and trained both arms from that same extraction with depth removed from one of them. The first seed favoured depth by +0.47. The second went the other way, −0.44. The third, +0.02. Averaged across seeds: +0.02 ± 0.46 — an effect twenty times smaller than the run-to-run noise. Depth resolves a real geometric ambiguity that 3,215-way word classification turns out not to need, because trajectory and gross handshape separate almost every pair long before fingertips matter.

Masking out low-confidence keypoints is standard hygiene. Does it help?

No — it was the most damaging thing we tried.

Pose extractors emit a confidence alongside each coordinate, and zeroing the ones below a threshold is a common and sensible-sounding step. It cost about 13 points of accuracy, across two seeds. The reason appears to be that hand keypoints carry low confidence most of the time — the hands are small in frame and often occlude each other — so a threshold that looks conservative deletes most of the hand most of the time. Whatever these confidence scores encode, it is not 'this coordinate is uninformative'. The model had been extracting usable signal from coordinates the extractor was not sure about.

Every modern translation system has a language model in the decoder. Does one help here?

No — it made things dramatically worse at our data scale.

Our translation model learned English from scratch on roughly twelve thousand sentence pairs, which is very little. The obvious fix is to hand the decoding to a pretrained language model and train the pose encoder to speak into it. We did that and BLEU fell from 8.9 to 0.8. The failure mode is legible in the output: the model produces fluent, confident English sentences that have nothing to do with the signing it was shown. With a weak visual signal and a strong language prior, learning to ignore the input is the easier way to reduce loss. We do not read this as evidence against language-model decoders in general — the field is right that they are the future — but as evidence that they need more grounded data than we currently have before they attend to the signer at all.

Can BSL data transfer to Auslan? They share a lineage.

Not through this route.

Auslan, BSL and NZSL form the BANZSL family with substantial vocabulary overlap, which makes cross-language transfer look attractive. Both encoder-transfer and co-training from BSL translation data produced null results on Auslan translation. A same-language control was also null, which locates the problem in the weak continuous-translation source rather than in the language gap.

Method

How we decide a number is real.

The depth experiment nearly produced a false headline. What stopped it is worth more than the experiment was.

Retraining the recogniser with depth restored, the first run came back +0.47 — a clear improvement, and the first positive representation change this project had produced. Reported on its own, it would have been wrong. The second run, identical in every respect but the random seed, came back −0.44.

Running both arms across three seeds gave the real picture: seed-to-seed variation on this benchmark spans up to 0.7 of a point, which is larger than most of the changes anyone would want to report. The depth effect itself was +0.02 ± 0.46 — indistinguishable from nothing.

So the rule is explicit: a difference smaller than roughly 0.7 points is not a result unless it survives multiple seeds, and comparisons are made between arms trained on the same data, on the same hardware, differing in one thing only. Where an older figure of ours does not clear that bar, we do not quote it as a difference.

None of this is novel methodology. It is standard practice that is easy to skip when a number comes out the way you hoped — which is exactly when skipping it costs the most.

Reading these numbers

Caveats that travel with every figure above.

  • Benchmark, not deployment. These are scores on a studio-collected dataset with clean clip boundaries. Continuous signing adds segmentation errors that compound with recognition errors.
  • Protocols differ between lines of work. Our recognition and translation tracks use different evaluation setups, so numbers from one should not be compared directly against the other.
  • Accuracy is not uniform across signs. Signs differing only in handshape or orientation are confused far more often than the headline suggests — see the hard-negative figure on the research page.
  • No Deaf-led evaluation yet. Every number here is a benchmark score, not an assessment by Auslan users of whether the output is useful or acceptable.