
Every language model provider caps how much traffic an account can send: requests per minute, tokens per minute, sometimes concurrent connections. Early prototypes never reach those caps. Production systems reach them on their busiest day, which is the worst possible moment to discover how the system behaves. Designing around LLM rate limits from the start is far cheaper than retrofitting it during an incident. Know your actual limits Limits vary by provider, model and account tier, and they change as usage grows. Find the current figures for the models you use and record them where the team can see them.
