Session

Offline RAG on Mobile (Flutter)

Most “AI in apps” quietly relies on cloud calls—slow, costly, and risky for PII. This session shows how to ship offline Retrieval-Augmented Generation in Flutter: chunk domain docs, create on-device embeddings, and run local k-NN to answer FAQs in <200 ms on mid-tier Android. We’ll compare vector storage (Isar vs SQLite), pick embedding routes (precomputed vs on-device via llama.cpp/GGUF or Gemini Nano where available), and wire a tiny local answerer (extractive or compact LLM). Live demo: Wi-Fi off, instant answers to policy questions.

Privacy jeeti, latency jeeti—aur cloud bill bole: shukriya bhai.

What attendees get

- A reference offline RAG pipeline (chunking → embeddings → local k-NN).

- Latency & size budgets to hit <200 ms P95 for ~5k chunks.

- A vector store recipe (schema, L2 normalization, isolates, k-tuning).

- A lightweight answering strategy (extractive vs tiny local LLM) with fallbacks.

- A minimal eval harness for latency, recall@k, and accuracy in CI.

Prashant R Dalai

Direction Software LLP, SDE

Mumbai, India

Actions

Please note that Sessionize is not responsible for the accuracy or validity of the data provided by speakers. If you suspect this profile to be fake or spam, please let us know.

Jump to top