Why Spatial Computing Should Learn From American Sign Language
I’ve studied input methods for most of my professional career, starting with keyboard-and-mouse combinations, then touch, and, most recently, voice interfaces that have failed to live up to their promise. Every transition has followed the same pattern: the industry spends years trying to force old interaction paradigms onto new hardware before realizing that the new hardware needs its own language.
We are at that point with spatial computing now. As far as I’m concerned, the answer has been staring us in the face for more than a hundred years.
I first encountered ASL in my second year of medical school, when I was studying to become a pharmacist. I left the program two years later to get into tech, but the ASL class I took has stayed with me ever since.
When Apple released Vision Pro and companies such as Meta and Samsung pushed their mixed-reality headsets into the market, they provided basic gesture sets for interacting with these interfaces: pinch to select, swipe to scroll, and pointing conventions borrowed from touch interfaces. Those gestures let users navigate, but anyone who has tried these devices understands their limits.
(Disclosure: Apple, Meta and Samsung are among the many global technology companies that subscribe to research from Creative Strategies, the firm I founded.)
A Vocabulary Problem, Not a Hardware Problem
A gesture can select an app, but today’s gesture vocabularies cannot easily convey degree, sequence, spatial relationship, or emotional inflection. The hardware is not necessarily the problem. The problem is vocabulary.
In 1997, Palm faced a version of this challenge with the PalmPilot . To use it effectively, people had to learn to write each letter in a form the device could recognize. Jeff Hawkins , the father of the PalmPilot, drove home that point when he showed me his invention before its release. He had previously worked at GRiD, an early laptop company that emphasized pen computing, and he understood that its pen computer lacked the power to interpret unrestricted handwriting. Instead, users had to follow specific stroke conventions so the system could recognize and digitize what they wrote.
Spatial computing risks making the same mistake: asking people to adapt themselves to a narrow input system rather than developing an interaction model suited to the medium.
What ASL Understands About Space
American Sign Language offers a compelling model. It is not spoken English translated into gestures. It is a complete natural language, with its own grammar and a long linguistic history. ASL conveys complex meaning through handshape, motion, location, and facial expression in three-dimensional space.
That distinction matters for spatial-computing designers . ASL did not evolve for a flat screen; it uses the signer’s surrounding space as part of the language. In ASL, a signer can assign a person, object, or idea to a specific spot in the space around them, then point or gesture toward that spot later to refer to it again. Researchers describe this use of signing space as a central way ASL expresses spatial relationships and maintains reference. That is strikingly similar to the persistent spatial anchors AR and VR designers are trying to create.
I have discussed this idea over the past few years with teams developing gesture systems at established headset manufacturers and smaller AR startups. Their reactions split in an interesting way.
Engineers working on low-level hand tracking often respond enthusiastically. They see a mature, well-documented visual language rather than a gesture set that must be invented from scratch. Research on handshape recognition and sign-language processing could also help inform future interaction systems, although recognizing fluent signing remains far more difficult than recognizing a handful of discrete commands. Robust systems must account for movement, handshape, body position, facial expression, variation among signers, lighting, and real-time performance.
The Difference Between Inspiration and Appropriation
However, product and user experience (UX) teams tend to react more cautiously, and rightly so. They worry that companies could appropriate the language of the Deaf community, pull out a few visually convenient signs, and reduce them to gesture shortcuts stripped of linguistic and cultural context.
Both reactions are valid. Reconciling them is the real opportunity.
The companies that get this right will not simply borrow a few ASL handshapes for product commands. They will work with Deaf linguists, Deaf designers, and the broader Deaf community to develop accessibility-first interaction models that respect ASL as a language. The goal should not be to make ASL a generic command vocabulary for hearing users. It should be to let ASL’s sophisticated use of space, expression, reference, and motion inform richer spatial interfaces—and to ensure that Deaf people can use those interfaces on equal terms.
Accessibility-First Design Benefits Everyone
This would not be the first time accessibility-first design improved life for everyone. Designers created curb cuts for wheelchair users, but they also made it easier for people to push strollers, roll suitcases, and navigate sidewalks. Closed captions began as an essential tool for Deaf people and now help hearing viewers follow videos when they cannot use sound.
Again and again, accessibility-driven design has produced better, more universal experiences. Spatial computing looks poised to become the next example.
There is also a hardware tailwind. The cameras and depth-sensing arrays needed to improve hand tracking overlap with the sensing capabilities that could support more sophisticated sign-language recognition. But the computer-vision challenge is not “almost entirely” solved. The more accurate claim is that the underlying sensing hardware is advancing quickly, while reliable interpretation of continuous, expressive, real-world signing remains a difficult technical problem.
What is still missing is the willingness to build a genuinely expressive interaction model on top of that hardware, rather than a poor imitation of one.
I do not expect a major platform to announce “ASL as the input method” anytime soon, and I am not sure that framing would be desirable. A more realistic path is gradual: expanded gesture vocabularies, better support for spatial reference and anchors, accessible communication features, and input models built collaboratively with Deaf experts from the beginning.
That is usually how the best ideas emerge in this industry: not through a big announcement, but by proving their value first.