BarunLM-35M converted to ONNX (fp32/fp16/int8) fully client-side with onnxruntime-web, PyTorch parity & fused GQA for WebGPU. - View it on GitHub
Star
3
Rank
3464480