Judgment Labs builds the continuous-improvement stack for AI agents: tooling to trace, evaluate, monitor, and improve agent behavior in production. Its open-source Judgeval SDK (Python and TypeScript, with Go and Java clients) instruments agent frameworks and model providers to capture traces, spans, and tool calls; Agent Judges and Code Judges… -
View it on GitHub