Session

The bill after the demo: what Gen AI really costs in production

Almost every company has shipped a Gen AI proof of concept by now. A growing number have received the first serious invoice.

This talk is the cost anatomy of a Gen AI application: what you actually pay for, and which architectural choices change the number by an order of magnitude.

We'll break down token economics and why prompt design is a cost decision before it's a quality decision. We'll explore when self hosting a model on Cloud Run with GPUs beats a managed API and, just as important, when it clearly doesn't. We'll see why scale to zero is the biggest lever most teams never touch, and how caching, batching and model routing (small model first, big model only when it earns it) reshape the bill.

The numbers come from projects whose invoices land in my own inbox: a multilingual AI podcast pipeline that publishes on a schedule, and a couple of MVPs running on Cloud Run. Small enough that I can show you every single line, real enough that the mistakes hurt. Everything else is normalised into cost per million tokens and ratios, so the lessons travel to any workload size, including yours.

You'll leave with a mental model to estimate Gen AI costs before you build, and a set of levers to pull when the invoice arrives anyway.


Your Gen AI proof of concept cost eleven euros. Your Gen AI product costs eleven thousand! Where does the money actually go? Token economics, Cloud Run with GPUs against a managed API, scale to zero, caching, batching and model routing. Real invoices from my own projects, small enough that I can show you every line. Come and learn to estimate the bill before you build it!

Nicola Guglielmi

GDE Cloud • Google Cloud Architect • Google Cloud Authorized Trainer • Team Manager • GDG Community Lead 🚀

Campobasso, Italy

Actions

Please note that Sessionize is not responsible for the accuracy or validity of the data provided by speakers. If you suspect this profile to be fake or spam, please let us know.

Jump to top