The Camera Behind The Skin
Execution·Framework·6 min read

The Camera Behind The Skin

A camera inside a soft gel dome reads texture at 10 to 50 micrometres, fine enough to read the ridges of a fingerprint, from a single sensor where an electrical array would need 10,000 taxels and 10,000 wires to see less. The cost is bulk, a 30 to 60 Hz bandwidth ceiling, and a frame that cannot thin below fingertip scale.

01

The Skin That Is Read From The Inside

A camera sits inside a dome of clear silicone and looks up at the back of it. The gel is 1 to 3 millimetres thick, coated on its inner face with a reflective layer, and the optics never see the object at all. They see the surface the object has just deformed. Press that dome onto a coin and the reflective layer changes shape. Press it onto a woven shirt and the thread pattern arrives as a picture. There is no grid of transducers underneath and no harness of wires running back to a board. One lens, one light, one cable. Johnson and Adelson published the first version in 2009. The company built on it was founded in 2014 as a spinout from MIT. That inversion is the entire idea, and it took the field twenty years to notice it. Touch did not need a better material. It needed to stop counting contact points and start photographing them.

02

Two Ways To Read The Same Dome

Two readout methods exist, and the difference between them decides what the sensor can feel. Intensity mode paints the inner surface with a reflective coating and reads brightness. Where the skin bulges the reflection brightens. Where it flattens, it darkens. The camera image is then a direct map of surface texture, at 10 to 50 micrometres, which is fine enough to read the ridges of a human fingerprint. Marker tracking prints a regular grid of dots on the same inner face. The camera follows how those dots move, and from that displacement field a full three-axis force distribution is reconstructed. Normal and shear separate cleanly. Effective pitch lands between 0.5 and 2 millimetres. The first method sees shape. The second feels load. OmniTact, built in 2020, does both from a hemispherical tip at roughly 0.3 millimetres. The sensor is not the gel. The sensor is the light coming back out of it.

03

One Cable Against Ten Thousand Wires

An electrical array that wants the same information needs roughly 10,000 taxels and 10,000 wires, and it still sees less, because every pixel of the camera contributes instead of one point per electrode. Routing problem gone. Power problem gone. The cable count drops from five figures to one. The second gift is the format. The output is an image, so every tool built for computer vision applies without translation. Convolutional networks, vision transformers and simulation-to-reality transfer all drop straight onto tactile data. Facebook AI Research open-sourced the DIGIT design in 2020, which put a working sensor into labs that had no sensor budget. GelSight Mini arrived in 2023 at 15 millimetres across. GelSlim solved the same problem in 2014 with a fingertip wedge and a pair of mirrors. One camera sees more of the contact than ten thousand wires ever will.

04

The Fifteen Millimetres You Cannot Remove

Then the bill arrives, and it is measured in millimetres. The camera needs focal distance. Between the gel and the lens sits 5 to 20 millimetres of empty space that light has to cross, plus internal illumination that has to stay uniform and stable. Total thickness cannot go below 10 to 15 millimetres, and that number does not negotiate. A sensor that cannot thin past a fingertip is a sensor for fingertips: patches, grips, tooling tips. Not a whole skin. Force numbers are respectable for a camera. Sensitivity reaches 0.01 newtons, and range runs 0.01 to 20 newtons, set by how compliant the gel is. It reads Braille dots spaced at half a millimetre. It is also a rigid camera sitting inside an otherwise soft structure, which puts every stress concentration at one seam. And the gel surface takes damage. Scratches and permanent deformation shift the calibration over weeks of ordinary use.

The sensor is not the gel. The sensor is the light coming back out of it.

05

Thirty Frames A Second Is The Ceiling

Bandwidth is set by the camera, not by the chemistry of the skin. Typical operation runs 30 to 60 frames per second, and a rolling shutter smears the fastest contact events into a blur that no post-processing recovers. Anything that happens inside a single frame interval is simply invisible. The compute line is the second cost. A neural network running on every frame needs a processor of its own, which is a strange thing to bolt onto a fingertip. There is a real probability that the next decade of progress in this field is spent chasing a frame rate it cannot raise, because the constraint sits in silicon that is already at its limit. You cannot make it thinner, and you cannot make it faster. You can only decide what it is for.

06

It Did Not Win On Performance

Nothing above is a winning specification. Bulk is worse than an electrical array. Latency is worse. The compute budget is worse. It dominates dexterous manipulation research anyway, and the reason is not a metric. The output is an image, and images arrived with a mature ecosystem attached. Contact rendered in simulation to train policies that transfer to hardware. Self-supervised learning methods that already existed for visual data. One shared encoder that can carry vision and touch together, because both arrive as the same kind of object. Ecosystem beats specification. It is the same pattern in every hardware market I have watched. I have built companies across twelve countries and deployed 210 energy systems, and the winner is almost never the best part. It is the part that everything else was already built to talk to. The probability that a team picks the sensor with the best datasheet and then stalls for two years on integration is not small. That is the ordinary outcome, not the exception. Touch spent forty years trying to build better skin. It won by borrowing eyes.

One camera sees more of the contact than ten thousand wires ever will.

The map is dead. Nobody told you.

Bali State of Mind is the survival guide for the collapse of everything you were taught to believe.

Beyond this book

Building the same thing somewhere else.

Julien Uhlig is available for advisory work, board seats and media appearances. Write to media@exventure.co.

The academy that trains the operators, across every company in the group, is EX Epic Academy - 25,000 applications, 25 seats per cohort, 210 alumni across 19 countries. academy.epicsolutiongroup.com

EX-AI Summit 2026

18-20 November. Online, Las Palmas, Bali.

Three days on what happens to work, capital and institutions when the map stops matching the ground. Seats are limited by cohort.

ex-aisummit.com →