Skip to content

Limitations

What this cannot do.

Stated first and in full, because a capability claim without its boundary is a misleading claim. Nothing on this site is validated for use where being wrong has consequences.

Do not use this for anything that matters

Not for medical, legal, emergency, financial, employment or educational assessment settings. Not as a substitute for a qualified Auslan interpreter. Not as evidence of what someone signed. The system has no way to know when it is wrong in the ways that would matter most in those settings.

Capability boundary

Where the research actually stops.

Isolated signs, not conversation

The recogniser names one sign from one clip. Connected signing — where signs blend, reshape each other and carry meaning through space — is a different and much harder problem that we have not solved.

A closed vocabulary

3,215 glosses is the model’s entire world. Fingerspelling, name signs, technical vocabulary and regional variants outside that list cannot be recognised — they will be silently mapped to the nearest thing the model does know.

Translation is barely working

Our continuous translation experiments score in the low BLEU range on everyday content and near zero on news. That is not a tuning problem — it is a data wall, and we report it as one.

Studio-shaped assumptions

Good lighting, a clear upper-body view, one signer facing the camera. Accuracy degrades outside those conditions, and the in-the-wild test set is the least flattering number we publish for exactly that reason.

73 signers of variation

The training data covers a limited range of signers. We have no basis for claiming the model works equally well across ages, regions, signing styles, or for signers whose bodies move differently.

No Deaf-led validation yet

Every number here is a benchmark score. None of it constitutes evaluation by Deaf Auslan users of whether the output is useful, accurate or acceptable — which is the evaluation that would actually count.

Interpreting the numbers

What an accuracy figure does and does not mean.

An accuracy of around 86% in the wild sounds close to working. It is not. It means roughly one sign in seven is wrong — and the model gives no reliable signal about which one.

Errors are also not uniformly distributed. Signs that differ only in handshape or orientation are confused far more often than the headline number suggests, which is why we publish the hard-negative scoring figure alongside it.

And benchmark accuracy is measured on clean clips with known boundaries. In continuous signing the system must also decide when a sign starts and ends, and those segmentation errors compound with recognition errors rather than averaging out.

Conditions

What would have to be true before this went further.

  • Deaf-led evaluation. Assessment by Deaf Auslan users of whether output is useful and acceptable, with the authority to say no.
  • Consent that covers the actual use. Especially for anything generative — synthesising signing from recordings of real people is a consent question before it is an engineering one.
  • Honest failure behaviour. A system that visibly declines when uncertain, rather than producing fluent output of unknown correctness.
  • A defined role alongside interpreters. Any deployment has to answer what happens when it is wrong, and who carries that.

Corrections welcome

If something here is wrong — a claim about Auslan, a mischaracterisation of Deaf community perspectives, or a number that does not hold up — we want to know. Corrections reach us through the Foundation’s get involved page.