
Evo-Depth
A lightweight vision-language-action model that uses implicit depth cues from RGB observations to improve spatially grounded manipulation without additional depth sensors.
My contribution
Model design, code implementation, training, and evaluation.
