
The Bottleneck Moved To The Software
In 2015 the limit on robotic touch was the sensor: too few taxels, too much drift, wiring that failed in the field. By 2025 an optical fingertip returns more than 100,000 effective taxels and an active-matrix skin more than 10,000. The hardware won. The interpretation did not, and the largest public tactile dataset holds about 100,000 samples against the billions that computer vision trains on.
Ten Thousand Fingertips And Nobody To Read Them
In 2015 the limiting factor in robotic touch was the sensor. Density was low, sensitivity was poor, and the wiring failed in the field. Every conference paper was a materials paper. By 2025 the numbers had flipped. An optical tactile sensor of the GelSight or DIGIT family returns more than 100,000 effective taxels from a single fingertip, because it reads deformation from a camera image rather than from a grid of discrete cells. Active-matrix electronic skins reach more than 10,000 taxels across a comparable area. We can now collect far more touch data than we can interpret. That is not a complaint about the hardware. It is the hardware having done its job so well that the constraint had to move somewhere else, and the somewhere else turned out to be software. A dense tactile stream is only worth what an algorithm can extract from it: this object is slipping, this is the edge of the part, this grasp is stable. The field did not solve the hard problem. It swapped it for a harder one.
Four Eras, Four Different Bottlenecks
Look at the last twenty-five years of electronic skin as a list of constraints rather than a list of inventions, and the pattern is almost too clean. From 2000 to 2015 the constraint was the transducer. The symptom was low sensitivity and high drift. The fix was better materials. From 2015 to 2022 the constraint moved to the interconnect. There were too few taxels and the data rate could not carry what existed. The fix was matrix arrays and active-matrix addressing. From 2022 to 2025 the constraint moved again, to interpretation. The symptom is a field that is data-rich and meaning-poor. The fix is models for tactile perception. And from 2025 onward the constraint is the sim-to-real gap. Models trained in simulation fail on real hardware, because contact simulation is nowhere near visual rendering. The fix has to be better contact simulators. Four eras, four symptoms, four fixes, and not one of them was the last one. Every fix handed the field a new bottleneck, downstream, and slightly more abstract than the one before. The materials people are done. The learning people are not.
A Hundred Thousand Samples Against A Billion
Computer vision trains on 10 to the 7 up to 10 to the 9 labelled samples, ImageNet at the low end, Common Crawl and its descendants at the high end. The largest public tactile dataset holds roughly 10 to the 5. Two to four orders of magnitude. That is the gap, and it is a data problem, not an architecture problem. There is no tactile ImageNet. There is no tactile GLUE, no standard benchmark that a research group can submit to and be ranked against its peers. Contact simulation sits far behind visual rendering in fidelity and in tooling. And the transfer problem is worse than in vision: a model trained on GelSight images does not run on a capacitive array, because the two sensors do not merely differ in resolution, they differ in what they physically measure. Anyone who has raced a system to production knows how this goes. Hardware improves on a curve you can buy. Datasets improve on a curve you have to organise, label, standardise and give away before anyone else benefits. One of those two curves is much easier to fund. That is why the sensor side looks solved and the software side looks stalled. The infrastructure is the product now. Nobody in the field is paid for building it.
Why The Camera Won
The optical sensors did not pull ahead because they are more sensitive. They pulled ahead because they produce an image, and an image arrives with an entire ecosystem already built around it. A tactile photograph can be fine-tuned into a pretrained ResNet or Vision Transformer, often touching only the last 2 or 3 layers. Data augmentation works on it in the standard way, rotation, scaling, brightness jitter, because it is an image and those operations are defined for images. Simulation engines built for robotics, PyBullet, Isaac Gym, MuJoCo, can render it. Self-supervised methods from the vision literature, MAE and SimCLR among them, transfer to it directly. The data format is the advantage. Optical fingertips return 1 to 10 megapixel RGB frames, and everything modern in machine learning already knows how to eat that. Now put a capacitive array beside it. It returns a 2D grid, typically 16 by 16 up to 100 by 100. A piezoelectric film returns a 1D time series. An event-based or neuromorphic sensor returns a sparse stream of addresses and timestamps that does not fit the standard frameworks at all. None of those are worse sensors. They are harder inputs, and the low channel count limits what features a network can find. A grid of 256 values has to learn a representation from scratch. A 1-megapixel image inherits one that cost a decade of someone else's compute to produce.
We solved the sensor. We did not solve the sentence.
What You Can Actually Ask A Fingertip
Strip the marketing out of tactile machine learning and it comes down to four families of question. Classification: what is this object, what is this surface, what is the contact doing right now, whether that is a stable grasp, a slip, a roll or a tap. Regression: how hard is the force, where is the object relative to the fingertip in position and orientation, and how close is the grasp to dropping. That last one, the slip margin, is the number a hand needs constantly and almost never reports. Representation learning: self-supervised embeddings learned from unlabelled touch, embeddings that transfer across sensors and tasks, and alignment between touch and the other two modalities a robot already has, vision and language. Control: policies learned from tactile feedback, model-predictive control that treats the tactile stream as its state estimate, and grip force adjustment running in real time off the sensor rather than off a schedule. The practice that has emerged around those four families is consistent. Start with an optical fingertip. Use pretrained vision backbones as feature extractors and fine-tune shallow. Fuse touch and vision early, at the feature level, by concatenation or cross-attention, never late by combining two decisions. And where the sensor is not an image, compress it with a learned representation first so the policy sees a compact feature rather than a raw stream. Fuse early. Almost every failure I have seen in instrumented systems came from letting two good measurements make their decisions separately and then trying to reconcile them afterwards.
The Missing Infrastructure Is Not A Genius Shortage
The probability that a tactile foundation model, one that transfers across sensor types and carries a benchmark the field agrees on, arrives within three years is not zero. Most road maps are pricing it at zero, because it is the one line item that no single company can own. What is missing is not a cleverer architecture. It is a dataset at 10 to the 8, a benchmark, a contact simulator that closes enough of the sim-to-real gap to be worth training against, and a common format so a model learned on one sensor can be asked something useful about another. Those are infrastructure. Infrastructure gets built when somebody decides that being early matters more than being first. I have built companies across twelve countries and restructured a 75 million euro industrial group, and the pattern is the same every time. The glamorous layer gets funded. The layer underneath it, the one everybody depends on and nobody owns, gets postponed until it becomes expensive. The hand can feel now. The machine still cannot say what it felt. That is the next decade of this work, and it is a software problem wearing a hardware budget.
One hundred thousand samples against a billion. That is the whole gap.
The map is dead. Nobody told you.
Bali State of Mind is the survival guide for the collapse of everything you were taught to believe.
Beyond this book
Building the same thing somewhere else.
Julien Uhlig is available for advisory work, board seats and media appearances. Write to media@exventure.co.
The academy that trains the operators, across every company in the group, is EX Epic Academy - 25,000 applications, 25 seats per cohort, 210 alumni across 19 countries. academy.epicsolutiongroup.com
18-20 November. Online, Las Palmas, Bali.
Three days on what happens to work, capital and institutions when the map stops matching the ground. Seats are limited by cohort.
ex-aisummit.com →