LionCity AI SalesOps | NUS-ISS hackathon team project
Our team built a prototype that moved WhatsApp sales enquiries through product lookup, pricing, stock and delivery checks, commercial approval and order creation. My work focused on integration and reliability: making the multi-turn workflow behave correctly, strengthening transaction handling and measuring the LLM usage behind the conversation.
The result was a working prototype with explicit business controls, alongside a clearer understanding of the work still required for production. It was not a customer rollout, and we did not measure revenue uplift or staff time saved.
The operational problem
A sales enquiry can arrive as a message like this:
"Same order as last month, but make it 300. Deliver Jurong Tuesday. Same price can?"
A salesperson has to turn that into a transaction. They need to identify the customer and products, retrieve the previous order, verify stock, establish the correct price, check delivery and decide whether the request needs approval.
We used that workflow as the basis for LionCity AI SalesOps. The aim was to let an AI agent coordinate routine steps without giving it authority to invent business facts or approve an exception.
What the team built
WhatsApp connected to a FastAPI application and an LLM-driven sales agent. Python tools handled customer lookup, product discovery, previous orders, inventory, pricing, delivery, date resolution, commercial authority and order creation. SQLite held application data. Streamlit provided the internal Sales Console, Business Data administration and SalesOps views.
The application was deployed to AWS with a GitHub-based staging workflow.

The authority boundary mattered more than the choice of model. The LLM interpreted the customer's language and selected tools. Tools supplied business facts. Python validated the business state and enforced configured rules. Humans decided exceptions.
A request for "100 industrial adapters with 10% discount" therefore needed a verified product, stock availability, applicable pricing and an authority check. A confident model response was not a substitute for those checks.
My contribution
This was a team project. I am not claiming that I individually built every component.
My contribution concentrated on the later integration and hardening work:
- Debugging multi-turn sales conversations and commercial approval flows.
- Investigating inconsistent totals and order-confirmation behaviour.
- Using regression tests to reproduce failures before fixing them.
- Strengthening the relationship between order persistence, inventory deduction and delivery-capacity consumption.
- Working on the AWS deployment workflow.
- Adding gateway diagnostics to investigate excessive LLM context and token consumption.
That work crossed the chat interface, API, model/tool loop, Python rules, database, tests and deployment. Most of the difficult bugs appeared between those layers.
Approval had to survive the conversation
Transactions within configured commercial limits could continue automatically. A transaction exceeding quantity, value or discount authority created a human approval request. After review, the conversation could continue using the approved terms.
A later customer reply might contain only "Yes please." The application still had to know which transaction that referred to, what had been verified, which terms were approved and whether an order already existed.
We increasingly separated natural-language history from trusted transaction state. History explained what had been said; validated state determined what the application could do next.
That distinction was necessary for both approval handling and order confirmation. It also exposed risks from stale state carrying into later conversations.
Reliability work after the happy path
Testing exposed faults that a convincing chat demo could conceal: a multi-item delivery summary used only the last item's subtotal, discounted orders showed pre-discount amounts at the wrong stage, and confirmation formatting was inconsistent. We also encountered risks in approval continuation, delivery-capacity consumption and stale transaction state.
I worked through these using a repeatable sequence: reproduce the fault, write a failing regression test, fix the application and run the broader suite.
One important change strengthened confirmed order creation so that order persistence, inventory deduction and delivery-capacity consumption behaved as one transactional operation rather than separate side effects.
This was a prototype hardening result, not proof of production resilience. The supplied write-up does not include a test-suite count or a production reliability measurement, so neither is claimed here.
Measuring the hidden model work
Near the end of the hackathon, the LLM quota was disappearing faster than expected. I instrumented the gateway rather than assuming that shortening a prompt would solve it.
A fresh "Hello" used 5,530 input tokens and generated 21 output tokens. In a controlled four-message sales conversation, the agent made nine LLM calls and used 60,490 input tokens.
The investigation showed two effects: one customer turn could trigger several model/tool iterations, and later calls carried accumulated tool traffic and conversation history. Those findings gave us concrete optimisation questions rather than a vague instruction to "use fewer tokens."
The companion article, Why four WhatsApp messages used 60,490 input tokens, contains the per-call measurements and the limits of that test.
Outcome and limits
The team demonstrated a WhatsApp-to-order prototype with verified business tools and a human approval path. My hardening work addressed transaction and multi-turn reliability, and the gateway investigation quantified model usage in one controlled sales flow.
The project does not establish sales conversion, labour savings or operating cost reductions. Those would require a deployed workflow, a baseline and repeatable measurement.
Before a production rollout, I would prioritise durable workflow state, authentication and authorisation, webhook verification, managed secrets, production-grade persistence, tracing, scalable deployment and bounded model context. Those are remaining requirements, not features I am claiming were delivered.
What this project demonstrates
The useful skill here was working through an ambiguous business process across several technical layers. I could reproduce a workflow failure, inspect the relevant state, add a regression test and explain why the fix belonged in application logic rather than in a more persuasive prompt.
That is the approach I would bring to an SME automation engagement: define which decisions can be automated, establish where the facts come from and make the exception path explicit before relying on an AI assistant.
Project card
LionCity AI SalesOps Applied AI · Python · FastAPI · WhatsApp · SQLite · Streamlit · AWS
Team-built hackathon prototype connecting WhatsApp enquiries to validated pricing, stock, delivery, approval and order workflows. My contribution focused on integration, regression testing, transactional correctness, deployment and LLM usage diagnostics.
Evidence note: project details and personal contribution are drawn from the supplied case-study draft and companion token-investigation draft. Source code, test reports, deployment records and diagnostic logs were not supplied for this editorial review.