Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

Considering moving from Groq Llama 3.3 70b to Gemini 2.5 Flash Lite for one of my use cases. Results are coming in great, and it's very fast (important for my real-time user perception needs).

What kind of rate limits do these new Gemini models have?



Are you using Groq Llama 3.3 70b from something like cline? Is it free and what are the API query limits?


I'm using it from their HTTP API. Limits I can't remember what they were initially tbh, I had to reach out through backchannels to get it increased to 300,000 tokens per minute.




Consider applying for YC's Fall 2026 batch! Applications are open till July 27.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: