We process about 2M tokens/day for our customer support RAG system. Switched from GPT-4o to Gemini 2.5 Flash and our monthly API bill went from $850 to $170. Quality is comparable — we ran a blind evaluation with our support team and they couldn't consistently tell which model generated which response. Flash is absurdly cost-effective for production workloads.
Workflows from the Neura Market marketplace related to this Gemini resource