We present a minimal system recipe for situated embodied conversation that pairs a real-time multimodal language model with a small set of tools for attention and active perception, enabling robots to interleave dialogue with “what to look at, when to look, and what to say” under tight latency constraints. - View it on GitHub
Star
0
Rank
14379828