PRODUCTION SCENARIO
A media archive needs alt-text and keyword tags generated once for 4 million photographs, with results due in three days. No user is waiting on any individual response.
Which processing option is the best fit?
Answering here is anonymous. Nothing is saved unless you sign in.
Show answer and explanation
Answer: Batch inference, which runs asynchronously at half the real-time price
Batch inference exists for large, non-urgent work: it is asynchronous and high-throughput, priced at a 50% discount to real-time inference, and most jobs finish within 24 hours. The real-time and reserved-capacity options cost more and buy latency guarantees that nobody here is waiting for.