Gitstar Ranking
Users
Organizations
Repositories
Rankings
Users
Organizations
Repositories
Sign in with GitHub
gmh5225
Fetched on 2026/07/13 21:13
gmh5225
/
pseudo_profiling_LLM
Tiny pseudo-profiling python script that estimates KV cache memory and a rough latency budget for sizing a deployment. (inputs: context length, target tokens, batch size, layers/heads/dim, dtype) -
View it on GitHub
Star
0
Rank
14124007