Negative results
Eight things that did not work.
Each cost real compute. Each is the kind of result that usually goes unpublished, which is precisely why it is here — if it saves another group the same spend, it did more good than a fifth accuracy point would have. They also have a shape: every one of them added information — depth, the face, appearance, another language, a language model. The intervention that did work added none, and changed where the cameras stood instead.
Does adding face landmarks to the classifier help?
No — it hurts.Adding 24 face points to the 55-point skeleton dropped top-1 from 84.7% to 83.6% on the same split. Facial grammar is real and it matters, but for isolated word recognition it added parameters without adding signal. Facial cues stayed in the system — they are routed to the interpretation layer as evidence, not folded into the classifier.
Does depth data make handshapes easier to tell apart?
No, not naively.The dataset is RGBD, so depth was free to try. Feeding it in alongside the skeleton made handshape classes less separable, not more. The plausible reading is that per-frame depth noise swamped the millimetre-scale differences that distinguish handshapes. Multi-view is the more promising direction.
Does a frozen RGB vision encoder add anything over pose?
No — it came out level.A dual-stream model with a frozen DINOv2 RGB branch scored 90.56% against 90.45% for pose alone on the same harness — inside noise. A handshape-specific probe showed the RGB branch was no better than pose at exactly the thing RGB was supposed to help with. The branch was dropped.
Fine — but does the face help continuous translation, where its grammar actually operates?
No. Same direction as before.The earlier face result came from isolated word recognition, where a single gloss gives the face no grammar to mark — so it was fair to say the test had been run in the wrong place. Questions, negation and topic marking exist only in connected signing. We rebuilt the translation data at 79 points and trained three seeds per arm, from scratch, on identical hardware. Adding the face moved BLEU from 8.87 to 8.48: a paired difference of −0.39, negative on all three seeds. We also checked the obvious objection — that the face signal was too small in scale to survive normalisation — and it was wrong: rescaling the face to its own frame doubled its spatial spread but left its motion over time unchanged. The likely limit is resolution. The dataset gives 68 face landmarks; a brow raise is a small deformation, and the face mesh running in your browser has 478.
The pipeline was throwing away depth. Does putting it back help?
No. And the way we found out is the point.The pipeline had been using two dimensions where the landmark model also provides a third. That omission is measurable: across 81,750 fingertip pairs, 57% of the pairs that look like contact in two dimensions are apart once depth is considered — and contact is one of the things that distinguishes handshapes. So we re-extracted the dataset keeping depth, and trained both arms from that same extraction with depth removed from one of them. The first seed favoured depth by +0.47. The second went the other way, −0.44. The third, +0.02. Averaged across seeds: +0.02 ± 0.46 — an effect twenty times smaller than the run-to-run noise. Depth resolves a real geometric ambiguity that 3,215-way word classification turns out not to need, because trajectory and gross handshape separate almost every pair long before fingertips matter.
Masking out low-confidence keypoints is standard hygiene. Does it help?
No — it was the most damaging thing we tried.Pose extractors emit a confidence alongside each coordinate, and zeroing the ones below a threshold is a common and sensible-sounding step. It cost about 13 points of accuracy, across two seeds. The reason appears to be that hand keypoints carry low confidence most of the time — the hands are small in frame and often occlude each other — so a threshold that looks conservative deletes most of the hand most of the time. Whatever these confidence scores encode, it is not 'this coordinate is uninformative'. The model had been extracting usable signal from coordinates the extractor was not sure about.
Every modern translation system has a language model in the decoder. Does one help here?
No — it made things dramatically worse at our data scale.Our translation model learned English from scratch on roughly twelve thousand sentence pairs, which is very little. The obvious fix is to hand the decoding to a pretrained language model and train the pose encoder to speak into it. We did that and BLEU fell from 8.9 to 0.8. The failure mode is legible in the output: the model produces fluent, confident English sentences that have nothing to do with the signing it was shown. With a weak visual signal and a strong language prior, learning to ignore the input is the easier way to reduce loss. We do not read this as evidence against language-model decoders in general — the field is right that they are the future — but as evidence that they need more grounded data than we currently have before they attend to the signer at all.
Can BSL data transfer to Auslan? They share a lineage.
Not through this route.Auslan, BSL and NZSL form the BANZSL family with substantial vocabulary overlap, which makes cross-language transfer look attractive. Both encoder-transfer and co-training from BSL translation data produced null results on Auslan translation. A same-language control was also null, which locates the problem in the weak continuous-translation source rather than in the language gap.