The language
Auslan is not English on the hands.
It is a full natural language, with its own grammar, its own spatial structure and its own culture. Every serious mistake in sign language technology starts by forgetting this.
Australian Sign Language is the language of the Australian Deaf community. It belongs to the BANZSL family, sharing a lineage with British and New Zealand Sign Language, and it is unrelated to American Sign Language despite both being used in English-speaking countries.
The 2021 Census recorded 16,242 Auslan users — the first time the language appeared as a Census prompt, following a long campaign by Deaf Australia. The community regards that figure as an undercount.
Auslan is not a manual encoding of English. It has its own word order, its own morphology, and grammatical machinery English simply does not have. Signs are inflected through space and movement: a verb can carry its subject and object in the direction it travels. Questions, conditionals, topics and negation are marked on the face and the body, running in parallel with the hands rather than in sequence.
Structure
Five parameters, all simultaneous.
A sign is not one gesture. It is five channels carrying meaning at the same time — which is why recognising signs is not the same shape of problem as recognising speech.
- HandshapeThe configuration of the fingers. Millimetre-scale differences separate distinct signs.
- OrientationWhere the palm and fingers point. The same handshape facing another way can be another sign entirely.
- LocationWhere the sign is made, relative to the body — and to places the signer has set up in space.
- MovementPath, direction, repetition and speed. Movement carries grammar, not just identity.
- Non-manual featuresEyebrows, eye gaze, mouth patterns, head and body position. These are grammar, not decoration.
Why it is hard
The gap between a word and a conversation.
What our models currently do is isolated recognition: given a clip containing one sign, name it. That is a well-defined problem with a benchmark and a leaderboard, and it is a long way from understanding signed language.
In real signing there are no gaps between words. Signs blend into one another, handshapes are pulled toward whatever comes next, and a signer sets up referents in the space around them and then points back to those locations for the rest of the conversation. Meaning lives in that spatial arrangement — and it is invisible to a model that only sees one clip at a time.
Then there is the part machines handle worst: depiction. Signers use their hands, body and face to show how something happened — its size, path, manner, and who was where. It is systematic and productive, and it does not reduce to a fixed vocabulary of signs to look up.
What this means for the research
Isolated recognition accuracy is a measure of a narrow capability, not of comprehension. We report it because it is what we can measure honestly — not because it is close to the thing that would actually help anyone.
Position
Technology has a bad track record here.
Sign language technology has a long history of projects that impressed hearing audiences and were useless — or offensive — to Deaf people. Signing gloves that captured handshape while ignoring facial grammar. Avatars deployed without Deaf consultation. Systems framed as fixing a deficit rather than removing a barrier.
The common failure is not technical. It is deciding what Deaf people need without asking them, and then measuring success by whether hearing people are impressed.
We are not exempt from that pattern by intention alone. It is the reason this site publishes negative results, states its limits plainly, and treats Deaf leadership as a precondition for validation rather than a consultation step at the end.