Run an Indian kirana store end-to-end from a Telegram chat — receive stock, cut GST-correct bills, run khata (credit), close the day, and pull PDF invoices & PPTX analysis decks, all in plain (Hindi/Hinglish/English) language. The chat is the product; the model orchestrates real tools over a consistent set of books.
Try it: spin up the Telegram bot yourself in one command (see Quickstart), or drive the agent locally with
scripts/sim_chat.py— no Telegram needed.
| Example message | |
|---|---|
| Receive stock | 50 packets of Maggi came in, cost ₹12, MRP ₹14 |
| Add a product | new item: Amul Butter 100g, GST 12%, MRP ₹62 |
| Cut a bill | make a bill: 2kg sugar, 1 Aashirvaad atta 5kg, 4 Maggi, 1 Amul butter, UPI |
| Edit mid-build | drop the butter, make it 6 Maggi |
| Stock / low-stock | how much sugar is left? · what's running out? |
| Khata (credit) | put ₹500 on Ramesh's credit · Ramesh paid ₹300 · Ramesh's balance? |
| Daily close | today's sales? / close the day |
| PDF invoice | send me that bill as a PDF → GST invoice document |
| Analysis deck | make this week's sales analysis deck → PPTX with real charts |
| Preferences | always assume UPI unless I say cash · default atta = Aashirvaad 5kg (remembered across /new) |
There is no menu and no intent router — every message goes through the model, which reasons over messy input, asks a clarifying question when genuinely ambiguous, and chains tool calls to get the job done.
# 0. deps (Python 3.11+, and Node for the Claude CLI the SDK drives)
uv sync # or: pip install -e ".[dev]"
npm install -g @anthropic-ai/claude-code # runtime dependency of claude-agent-sdk
# 1. secrets
cp .env.example .env # add ANTHROPIC_API_KEY and TELEGRAM_BOT_TOKEN
# 2. seed a realistic kirana catalog (+ optional sales history)
uv run python scripts/seed.py --reset --with-history
# 3a. run the Telegram bot (long polling — no public URL needed)
uv run python -m kirana.telegram.bot
# 3b. or drive the agent locally without Telegram
uv run python scripts/sim_chat.py
# 3c. or run the full scripted end-to-end demo (this is the recording script)
uv run python scripts/demo.pyDeploy: it's a long-polling worker (no inbound port or domain needed), so docker compose up -d --build runs it on any Docker host; a render.yaml is also included. The Dockerfile installs Node + the Claude CLI, and the compose file mounts a volume so the SQLite store, khata and invoices survive redeploys.
These concepts — skills, tool schemas, subagents, memory — map 1:1 onto the Claude Agent SDK, so authoring the capability surface is idiomatic rather than bolted-on:
- In-process tools via
@tool+create_sdk_mcp_server(no separate service, no network hop) — 30 thin tools exposed asmcp__kirana__*. - Filesystem skills (
.claude/skills/*/SKILL.md) loaded withsetting_sources=["project"]+skills="all"— the model pulls in the right domain playbook on demand. - PreToolUse hooks for a guardrail/audit backstop.
- Sessions with
resume=for multi-turn context that survives restarts. - Python also has the best libraries for the two artifact deliverables — reportlab (GST invoice PDF) and python-pptx (deck with native, editable charts).
Runtime model defaults to Sonnet 5 (excellent tool-orchestration, cost/latency-friendly for a well-scoped tool surface); set KIRANA_MODEL=claude-opus-4-8 for more headroom.
Telegram msg ──▶ dedupe(update_id) ──▶ session.run_turn(chat_id, text)
│ query(prompt, resume=session_id)
▼
Claude Agent SDK loop: observe → reason → act → observe …
│ calls mcp__kirana__* (multiple per turn)
▼
tools/*.py (thin) ──▶ db/repository.py (transactional) ──▶ SQLite (WAL)
│
final reply + queued PDF/PPTX ──▶ sent back to the chat
One inbound message = one agent turn that may chain many tool calls. The chat maps to a persisted Claude session (chat_id → session_id in SQLite) so context survives process restarts; /new drops the session but never the store.
Skills route; tools are thin adapters; the repository owns every business rule. A tool never does GST math or checks stock — it parses args (rupees→paise, qty+unit), calls one repository method, and returns a small model-friendly result.
| Skill | Tools (mcp__kirana__…) |
|---|---|
receiving-stock |
add_product, receive_stock, search_products, get_stock, list_low_stock, reorder_suggestions, list_expiring |
billing |
cart_add_item, cart_update_item, cart_remove_item, cart_view, cart_set_payment, cart_set_customer, cart_finalize, cart_cancel |
khata-credit |
khata_add_credit, khata_record_payment, khata_balance, khata_statement, list_khata_debtors |
store-analytics |
daily_summary, sales_report, top_items, gst_collected, stock_health |
documents |
generate_invoice_pdf, generate_analysis_deck |
store-memory |
set_preference, get_preferences, set_shop_profile, get_shop_profile |
- Grounding — prices/GST/stock only ever come from tool results; the system prompt forbids inventing SKUs, and every bill starts with
search_products. - Oversell guard —
finalizerunsUPDATE products SET stock_base = stock_base - :q WHERE id=:id AND stock_base >= :qand checksrowcount. Enforced in SQL, not the prompt — a jailbroken model still cannot oversell. - GST correctness —
domain/gst.py: per-item slab, back-calculated taxable from the GST-inclusive shelf price, CGST=SGST split that always reconciles to the paise, rupee round-off, legible per-slab breakup. All money is integer paise; quantities are integer base-units (no floats). - Multi-turn bills — the draft cart is a first-class
billsrow (statusdraft), so edits persist and a restart mid-bill loses nothing. Stock decrements only onfinalize. - Idempotency —
processed_updates(update_id)drops Telegram redeliveries;bills.idempotency_key UNIQUEmeans a retried finalize returns the same invoice with no double-decrement (proven under concurrent duplicate finalizes). - Concurrency — WAL +
BEGIN IMMEDIATE+ a write lock + the conditional decrement. Test: 25 buyers race for 10 packets → exactly 10 succeed, stock lands at 0, never negative. - Guardrails — no hard-delete tool; selling below cost or over-settling a khata returns a confirm-or-refuse error the model relays; khata to an unknown customer is refused.
- Artifacts — reportlab → branded, GST-correct PDF (HSN, CGST/SGST, amount-in-words, ₹ glyph); python-pptx → 8-slide deck with 4 native charts (top items, payment mix, daily trend, GST by slab) + computed insights.
- Memory across sessions —
preferences+shop_profilelive in SQLite and are injected into the system prompt at session-build time./newclears the conversation, not the store — defaults (payment, brand, GSTIN) still apply. - Expiry / FEFO — each stock-in is a
batchesrow with an optional expiry; sales draw down the earliest-expiry batch first,SUM(batches)stays equal tostock_base(the oversell guard is untouched), andlist_expiringsurfaces what's about to lapse.
- Branded/templated invoice PDF — shop identity header, per-slab GST, amount-in-words.
- Reorder suggestions from sales velocity — days-of-cover from recent movement.
- Multi-language — Hindi / Hinglish understanding and replies.
- Expiry / batch tracking with FEFO — dated batches, earliest-expiry-first consumption,
list_expiringalerts. - Scheduled deck + khata reminders — a JobQueue auto-sends a weekly analysis deck and daily khata-outstanding reminders to the owner (no LLM cost; configurable via
KIRANA_*env).
src/kirana/
domain/ money.py · units.py · gst.py · models.py # pure logic, no I/O
db/ schema.sql · connection.py · repository.py # transactional store (all invariants)
tools/ inventory·billing·khata·analytics·documents·preferences + registry, context, _helpers
artifacts/ invoice_pdf.py · analysis_deck.py
agent/ system_prompt.py · options.py · guardrail_hooks.py · session.py # the control loop
telegram/ bot.py # long-polling adapter
.claude/skills/<name>/SKILL.md # the routing playbooks
scripts/ seed.py · demo.py · sim_chat.py
tests/ 51 tests — GST, units, money, oversell, idempotency, concurrency, khata, artifacts, tools
uv run pytest -q # 51 tests, no API key needed — proves every hard part deterministically
uv run python scripts/demo.py # drives the REAL agent through all scenarios with DB checkpointsRead the case study → — the design decisions, the trade-offs, and what broke along the way.
Built by Harsh Kedia.