Skip to content

Dataset

What the models learned from.

Everything here is trained on MM-WLAuslan. Its licence is non-commercial, and that single fact shapes what this initiative is allowed to become.

MM-WLAuslan

A word-level Auslan corpus, recorded from three views.

Published at NeurIPS 2024 as a datasets and benchmarks track paper by the University of Queensland and collaborators. It is, as far as we know, the largest word-level Auslan recognition dataset that exists.

282KSign videosRecorded in a studio environment
3,215Auslan glossesCommonly used vocabulary — the model’s entire world
73SignersThe ceiling on how much signer variation we can learn
4Cameras, RGBDThree Kinect-V2 plus one RealSense, triple-view

Source: Shen et al., MM-WLAuslan: Multi-View Multi-Modal Word-Level Australian Sign Language Recognition Dataset, NeurIPS 2024 Datasets and Benchmarks Track (arXiv:2410.19488).

Test design

Four test conditions, not one.

The dataset ships four separate test subsets, each covering the full vocabulary. This is the most useful thing about it: it makes robustness measurable instead of assumed.

STU — studio

Same conditions as training. The friendliest number a model will ever produce, and the one to be most suspicious of.

TED — temporal disturbance

Timing is altered. Tests whether the model learned the movement or just the tempo it was trained on.

SYN — synthetic background

Background replaced. Tests whether the model is reading the signer or reading the room.

ITW — in the wild

Real-world capture conditions. The closest thing to an honest estimate, and the number we lead with.

Licence

CC BY-NC-SA 4.0 — and we mean the NC.

MM-WLAuslan is released under Creative Commons Attribution-NonCommercial-ShareAlike 4.0. Anything trained on it inherits that constraint. Our models are therefore non-commercial research artefacts, and this initiative sits inside the Octopus Foundation rather than inside a company.

This is a real limit, not a formality. A commercial Auslan product could not be built on these weights. It would require either purpose-collected data with consent obtained for commercial use, or a separate licensing agreement with the dataset’s authors. We have deliberately not blurred that line.

Consent is not a licence checkbox

The recordings are of real people signing. A permissive licence settles what we may legally do; it does not settle what those signers understood they were agreeing to. Any future use that would generate synthetic signing from those recordings — an avatar, for instance — is a consent question first and an engineering question second.

What the data cannot tell us

A studio corpus has a studio’s blind spots.

  • Isolated signs only. One sign per clip, with clean boundaries. Real signing has no gaps, and signs reshape each other where they meet.
  • 73 signers is not a population. Regional variation, age variation, and the natural range of signing styles are all under-sampled relative to the real community.
  • Vocabulary is a closed world. The model can only ever name one of 3,215 glosses. Fingerspelling, name signs, and depicting constructions fall outside it entirely.
  • Studio framing. Even the in-the-wild subset was built by a research team, not scraped from how people actually use video.