All posts

How You Touch Matters: Action-Conditioned Tactile Affordance for Robots

Press your thumb slowly into a kitchen sponge and it gives way smoothly, then springs back. Drag the same thumb quickly across it and something else happens: the surface catches, shears, maybe tears a little. Do the same on a tabletop and pressing or sliding feels almost the same, firm and unchanging. The place is the same and the hand is the same. Only the action differs, and you feel something different.

People use this all the time without thinking about it. We test a step with some weight before trusting it, or drag a finger across a surface to see whether a box will slide on it. Most robots do not reason this way. This post is about FRACTAL, the model at the center of my PhD, which tries to give them that ability. The paper is under review at ICRA 2027, so I describe the idea and the system here and leave the numbers for after the review.

The problem: maps that answer "where", not "how"

Robots usually use tactile sensing reactively and locally: detect slip, classify a texture, check that a grasp holds. More ambitious systems build tactile maps. Tactile SLAM, for example, can recover object geometry and pose with good accuracy. That tells a robot where things are and what they are.

When robots do reason about what they can do with a surface, the answer is usually stored on a static map: a label, a suitability score, or a region picked out by a vision model. Each location gets one affordance, whatever the robot plans to do there.

The sponge shows why that breaks. A compliant surface can yield elastically under a slow normal press and slip or yield plastically under a faster tangential probe at the same spot. A single label per location cannot capture this, because the outcome belongs to the pair (location, action), not to the location.

The idea: separate what lasts from what happens

FRACTAL is short for a Generative Factorized model of Action-Conditioned Tactile Affordance. The central move is to split the world into two parts.

The first is a physical state field, which does not change when you touch it: where the surface is, and its material constants such as effective stiffness, yield pressure and friction. The robot never observes this field directly, so it keeps a belief over it, a Gaussian process that also records where it is still uncertain.

The second is the interaction regime: the character of a particular contact while it happens. No contact, elastic loading, plastic yield, kinetic sliding, or combinations such as sliding while yielding. The regime is transient and depends on the action. It also has its own dynamics. During one sustained contact, static friction comes before slip, for example.

The sensor readings are then treated as exhaust, a noisy, sensor-specific trace of the regime. The model does not try to explain every taxel value. It asks which physical event most likely produced those values. Because the observation model is generative, the robot can also run it forward and imagine what a touch would feel like before it makes that touch.

Affordance becomes a prediction

With that structure, an affordance is no longer something you look up. It is a question you ask: if I perform this action here, what regime will I probably cause, and is that regime good for my goal?

FRACTAL answers by simulating the observation the action would produce, updating its belief about the regime from that imagined evidence, and scoring the result against the task. I studied two goals. Support looks for a flat surface that does not yield, where something can rest. Transport looks for a steady slope where sliding is easy to induce. A sponge-like patch may score badly for support under a firm press and well for transport under a slide, which is exactly the distinction a fixed label cannot make.

Choosing the next touch: reward vs information

A robot exploring by touch has a limited budget. Every probe costs time and effort, so the choice of the next probe matters.

The prediction above gives two things at once. It gives the expected reward of an action, which is how useful the action is for the goal. It also gives, for free, the expected information gain: how much the action would sharpen the robot's belief about the regime. FRACTAL combines the two into a single Expected Free Energy objective, adds a small penalty on effort, and picks the action with the best score.

This produces a natural pattern of behavior. At the start, every candidate looks equally (un)promising, so the information term dominates and the robot explores, probing where it is most unsure. As the belief sharpens, the reward term takes over and the robot exploits. The switch from exploring to exploiting needs no separate schedule or second weight to tune. It comes from the objective itself.

In simulation I compared FRACTAL with an action-conditioned baseline that predicts outcomes from location and action but has no regime inference, no uncertainty over the regime and no memory across contacts. Inferring the regime explicitly improved task success and reduced regret for both goals. The figures will follow once the review is over.

Loop closure: recognizing a place by how it feels

Tactile exploration has a familiar mapping problem. Small errors in the robot's pose estimate pile up into drift. Vision-based SLAM corrects drift by recognizing places it has seen before. Touch can do the same, as long as you have a good enough "fingerprint" of a place.

FRACTAL's fingerprint combines geometry, material properties and the full uncertainty over the interaction regime, not just its most likely value. Two spots match only when they have the same shape, the same material and the same ambiguity about how they respond. In a simulated pose-graph experiment, this physically informed descriptor gave a clear improvement in drift correction. For now, that experiment takes the relative pose of a recognized place from the simulator, and building the tactile registration step that would supply it is still open.

The hardware: two ways of feeling the same contact

A Franka Panda arm probing a textured white terrain map, and a close-up of the end-effector with its compliant four-bar finger The Franka Panda over one of the terrain maps (left) and the dual-channel tactile end-effector (right).

On the real robot, a 7-DoF Franka Emika Panda carries a custom end-effector with two tactile channels. The main surface is a PaXini GEN3 Omega L5321 array with 239 three-axis taxels over 53.38 × 25.10 mm. Beside it is a compliant finger: a four-bar linkage driven by a small motor, with a PaXini GEN3 S1813 Elite fingertip sensor of 31 taxels. Both channels see the same contact through different mechanical paths, and that is what allows FRACTAL's cross-modal consistency check.

MoveIt 2 plans the free-space motion only. Near the surface, a joint-space impedance controller takes over and runs a four-phase probe: approach, contact acquisition, interaction (press, slide or tap) and retraction. I ran qualitative case studies on two terrain maps, one in foam and one 3D-printed, to show that the whole loop works on real, noisy, multi-channel data.

Limitations and what is next

I want to be clear about scope. The simulator is a deliberate 2D, single-channel reduction, so the cross-modal factor and the full 3D submap graph were not tested there. The hardware runs are case studies, not a statistically powered comparison. The paper names the next steps: a full multi-modal 3D system, hardware ablations of the regime's contribution, repeated-trial comparisons on the robot, and isolating the cross-modal factor on real data.

The broader point is simple. For touch, "where" is only half the question. A robot that knows how a surface will respond to how it is touched can choose its actions, and its next probe, much more intelligently.

Read more