Hardware-in-the-loop benchmark for agentic embedded AI deployment. Agents iteratively edit model or firmware artifacts, the framework compiles and flashes them, and the score comes from real hardware measurements such as deployability, current, energy, and temperature. -
View it on GitHub