Context
Imitation learning for fine bimanual tasks (ziploc, slotting a battery, velcro) on a cheap rig. 1) ALOHA - teleop setup, 2) ACT - policy to predict the next actions from one observation instead of just the next one
A lot of stuff like CVAE, temporal ensembling isn’t relevant anymore but chunking is in pretty much every policy since (Diffusion Policy, GR00T, π), and ACT is a popular first policy
ALOHA
- <$20k total. 2 ViperX 6-DoF follower arms + 2 smaller WidowX leader arms, human backdrives the leaders and the followers copy joint for joint
- Joint-space teleop instead of VR/spacemouse no IK, so no singularities and it’s way more intuitive for fast precise stuff
- 4 webcams (2 wrist, 1 front, 1 top) at 480×640, everything runs at 50Hz
- 3D-printed see-through fingers + grip tape. Arms are only 5-8mm accurate so the policy has to close the loop visually
- User study: 6 people teleoping at 5Hz vs 50Hz, ~62% slower at 5Hz (p<0.001). Their argument for why fine manipulation needs high-freq control
Pipeline
- Obs = 4 RGB images + 14 joint positions (7+7). Action = absolute target joint positions, also 14-dim. Policy outputs at once
- Each image + 2D sinusoidal pos embeddings. 4 cams
- Append joint positions + style variable (both projected to 512) into the transformer encoder (4 layers)
- DETR-style decoder (7 layers): fixed position embeddings as queries cross-attend to the encoder output
- Trained as a CVAE. Encoder (BERT-ish, training only) reads [CLS] + joints + the demo action chunk (no images, “for faster training”) and outputs for a 32-dim
- Loss with . L1 over L2 bc it was more precise
- Inference: drop the encoder, set (prior mean)
- ~80M params, ~5h on a 2080 Ti, ~0.01s per inference
- Data: 50 demos per task (100 for velcro) = 10-20 min of demos, 30-60 min wall clock w/ resets
Temporal ensembling: query the policy every step so step has up to overlapping predictions, then average them with where is the oldest. Smaller = new obs gets folded in faster. Note older predictions get weighted more, which feels backwards but it’s for smoothness
Baseline
BC-ConvMLP, BeT, RT-1, VINN (all single-step, BeT/RT-1 use history + discretized actions)
Results
TABLE I: Success rate (%) for 2 simulated and 2 real-world tasks, comparing our method with 4 baselines. For the two simulated tasks, we report [training with scripted data | training with human data], with 3 seeds and 50 policy evaluations each. For the real-world tasks, we report training with human data, with 1 seed and 25 evaluations. Overall, ACT significantly outperforms previous methods.
TABLE II: Success rate (%) for the remaining 3 real-world tasks. We only compare with the best performing baseline BeT.
- Velcro roughly halves every stage (92% 20%): gripper closes too early on the mid-air grab, insertion not precise enough
- Human data is a lot harder than scripted for every method (multimodal + stochastic)
- Failures: candy unwrapping 0% (can’t find the seam on a low contrast wrapper), small ziploc mid-air. Hardware can’t do high force or anything that needs fingernails
What held up
Comparing the paper’s ablations with what people found later.
Chunking is the part that matters. Going from to takes success from 1% to 44% (open loop, no TE, averaged over the sim tasks), and bolting chunking onto BC-ConvMLP and VINN helps them too, so it isn’t tied to the ACT architecture.
Temporal ensembling gave +3.3% in the paper, but LeRobot turns it off by default and it’s been reported to hurt on PushT.
The CVAE is shakier. The paper says removing it drops human-data success from 35.3% to 2%, but Tony Zhao said in Sep 2023 that “the CVAE might not matter much”. On PushT runs has been found to carry nothing (mean within 0.005 of 0 in all 32 dims) and turning it off makes no difference, and a 2026 preprint finds the same empty latent on ACT’s own sim tasks.
The paper kinda skips how many of the 100 actions to actually execute before replanning (n_action_steps in LeRobot), and it matters a lot. On PushT at 10Hz, executing all 100 got 0.9% and executing 16 got 20%. Probably ~2s of commitment is about right for this kind of task: 100 steps at 50Hz is 2s, but 100 at 10Hz is 10s. So chunk length should be thought of in seconds, not steps.
Chunking vs one-step BC
Demos aren’t Markovian. A demonstrator hovers before grabbing, so for that image + joint state the dataset has ~15 frames of “stay still” and 1 of “move”. A one-step policy learns to stay still, and since nothing changes the frame it sits there forever (the paper’s robot “can pause indefinitely for certain states”). A chunk predicted during the hover contains the wait and the move, and executing it commits to the move.
Why chunking helps with compounding errors is not actually settled. The paper’s argument is that actions per decision means fewer decisions for errors to compound over. Zhang et al. say the arm + servo are stable on their own so deviations die out anyway. Lazzati et al. say a chunk acts on an older obs, which is more in-distribution. Zeng et al. argue compounding error isn’t the real problem and the pauses matter more.
CVAE vs plain regression
Regression averages demos that disagree. With L2 you get the mean, so if half the demos go left of an obstacle and half go right, the mean goes through it. Even a single route breaks it if the demos run at different speeds, bc at a given offset into the chunk they’re spread along the route. L1 (what ACT uses) gives the per-DoF median instead, which can stitch a chunk together from different demos per DoF, so it’s still OOD.
The CVAE gives the decoder an extra input that picks which demo to follow. The obs goes straight into the decoder untaxed and only pays the KL tax, so should only encode what the obs can’t explain (left vs right, fast vs slow). At inference , which only gets you one demo’s chunk if the encoder actually used and the demos near 0 agree. In practice it often collapses and you’re back to L1 regression + chunking. Most later work swapped the CVAE for diffusion / flow matching, which actually models the multimodal distribution Diffusion Policy
Future directions
Chunk boundaries vs reactivity is still the main tension. TE was ACT’s fix; PI’s real-time chunking (RTC) instead inpaints the next chunk while the current one is still executing. OpenVLA was single-step, and OpenVLA-OFT later added chunking + L1 regression and got a big jump.
On the hardware side ALOHA led to Mobile ALOHA and ALOHA 2, and leader-follower joint-space teleop is now the standard way to collect bimanual data.