
A robot may eventually be able to handle routine 3D printing chores simply because you told it what to do.
Researchers from several US universities have demonstrated ChunkVLA-AM, a system that combines a vision-language-action model with a FAIRINO FR3 robot for post-print handling. The idea is to let the robot look at the workcell, understand a short language instruction and generate the motions needed to carry it out.
That sounds a bit like attaching ChatGPT to a robot arm, although the underlying system is far more specialized. The researchers used OpenVLA-OFT, a seven-billion-parameter vision-language-action model, and adapted it to the particular motions and coordinate system of the FR3 robot.
Instead of asking the AI model for one tiny motion, executing it, taking another image and asking again, ChunkVLA-AM predicts a sequence of eight actions in one request.
The robot then performs that short sequence before taking a fresh camera image and requesting the next eight. This gives the system smoother motion without leaving the robot completely open loop for very long.
Each action includes changes in X, Y and Z position, three rotation values and a gripper command. The robot-side software also checks proposed motions before executing them, rejecting commands that leave the permitted workspace, move too far, violate minimum clearance or arrive after an emergency stop.
There is another interesting architectural detail: the AI model does not run on the robot computer. The researchers ran the 7B model on a remote server equipped with an RTX 6000 Ada GPU. The robot workstation itself remained CPU-only, sending camera images to the server and receiving action chunks in return.
For the tests, the system used only a single RGB camera. There was no depth camera, tactile sensing or multiple-camera arrangement.
The researchers first adapted the model to the FR3 using recorded demonstrations and a parameter-efficient training technique called LoRA. That was essential. Zero-shot versions of the models produced position errors measured in hundreds of millimeters, making them unusable in the small workcell.
With adaptation and eight-step chunks, the reported average position error dropped to 1.74mm. The X and Y errors were below 1mm, although the Z-axis error remained considerably larger at 3.87mm.
The physical test was deliberately simple. The robot was instructed to “catch the blue block” or “catch the red block,” pick the requested object from point A and move it to point B.
That is only a proxy for retrieving a completed 3D printed part, but it tests something conventional industrial robot programs usually avoid: choosing an object based on both what the camera sees and what an operator says.
Across 42 physical trials, the robot completed 39 transfers, a 92.9% success rate. The three failures all occurred during placement, where poor release-height control caused the object to fall and topple.
Lighting also turned out to matter. The researchers found the lowest prediction errors at intermediate brightness levels, while very dark and strongly overexposed scenes produced larger errors. That could become important around reflective metal build plates, dark polymer parts and changing machine illumination.
A more capable version could potentially be told to retrieve a finished part, inspect a warped edge or remove a failed print without someone first programming every motion. Those more complex AM tasks are among the areas the researchers identify for future work.
That would make the robot less like another piece of fixed automation and a little more like an operator that can see the problem you’re talking about.
Via arXiv
