PRODUCTION SCENARIO
A telecom routes 6 million short support messages a day into 12 categories. Latency must stay under one second and the per-message budget is a fraction of a cent.
Which Gemini model tier is the most appropriate starting choice?
Answering here is anonymous. Nothing is saved unless you sign in.
Show answer and explanation
Answer: A Flash-Lite model, the most cost-efficient tier for high-volume, low-latency work
Flash-Lite is Google's most cost-efficient Gemini tier, tuned for low-latency, high-volume, cost-sensitive traffic, which is what short-message classification at this scale is. Pro models are positioned for complex reasoning and coding, and Live API models target real-time voice and video rather than bulk text classification.