By Syahmul Aziz | LionCity AI SalesOps | NUS-ISS hackathon
A controlled test of our WhatsApp sales-agent prototype produced nine LLM calls and 60,490 input tokens from four customer messages. A fresh conversation containing only "Hello" already used 5,530 input tokens.
I investigated the usage by adding diagnostics to the LLM gateway. The useful finding was that visible conversation length was a poor guide to model consumption: the application made repeated calls, and those calls carried a growing context.
The quota warning
Near the end of our NUS-ISS hackathon, the team's allocated LLM quota was disappearing faster than the WhatsApp message volume suggested. A teammate suspected that the application was retaining conversation history and resending it on every request.
Our prototype, LionCity AI SalesOps, connected WhatsApp to a FastAPI application, an LLM sales agent, Python business tools, SQLite and a Streamlit console. The agent could look up products, check stock and pricing, resolve delivery requirements, request commercial approval and create an order.
That meant a short customer message could start several steps inside the application. Before changing the prompts, I wanted to count those steps and measure what each request carried.
Measure at the gateway
I added lightweight observability to the gateway adapter rather than changing the agent's behaviour. The diagnostic captured gateway-reported input and output tokens, message count, tool count and serialized payload size. It did not log customer message content or API secrets.
The first baseline used a fresh conversation:
Customer: "Hello" Gateway report: 5,530 input tokens and 21 output tokens.
The greeting contained almost no text. The request still carried system instructions, tool definitions and the message wrapper. That established a substantial starting overhead; it did not isolate the exact contribution of each component.
Four customer messages, nine model calls
I then ran a short sales flow:
- "Hello"
- "I want to buy 5 Industrial Cable and 5 Industrial Adapter."
- "Deliver to Tengah tomorrow."
- "Yes please."
Those four turns generated the following gateway measurements:
| LLM call | Messages in request | Serialized payload | Input tokens | Output tokens | Result |
|---|---|---|---|---|---|
| 1 | 2 | 21.1 KB | 5,530 | 21 | Final response |
| 2 | 4 | 21.4 KB | 5,587 | 219 | Tool use |
| 3 | 9 | 22.9 KB | 6,180 | 299 | Tool use |
| 4 | 14 | 24.3 KB | 6,835 | 98 | Final response |
| 5 | 16 | 24.7 KB | 6,960 | 79 | Tool use |
| 6 | 18 | 25.0 KB | 7,079 | 81 | Tool use |
| 7 | 20 | 25.5 KB | 7,240 | 111 | Final response |
| 8 | 22 | 25.9 KB | 7,387 | 137 | Tool use |
| 9 | 24 | 26.7 KB | 7,692 | 278 | Order creation |
The totals were 60,490 input tokens and 1,323 output tokens: 61,813 tokens processed. Input accounted for about 97.9% of that total. The application made an average of 2.25 LLM calls per customer turn in this test.

Two effects multiplied the usage
The first was repeated model invocation. The agent could request a tool, receive its result, call the model again and request another tool before answering the customer. One WhatsApp turn was not necessarily one LLM call.
The second was context accumulation. Assistant tool requests and their JSON results were appended to conversation history. Later calls carried that history alongside the system instructions and tool definitions.
By the ninth call, the request contained 24 messages rather than two. Input tokens had risen from 5,530 to 7,692, an increase of about 39.1%.

A better way to estimate usage is:
Total input tokens = the sum of the context sent in every model call.
Counting customer messages alone misses both the extra calls and the larger payloads. Serialized payload size is useful diagnostic information, but it is not itself a token count.
What the test does and does not prove
These are measurements from one prototype configuration and one controlled conversation. They are not a universal cost for WhatsApp or agentic AI.
The organisers reportedly saw roughly 200 large requests from our team. My test showed how substantial usage could accumulate, but it did not reconcile the entire quota against a complete request log.
It also did not establish a monetary bill. Token totals, provider pricing, prompt-cache treatment and actual billing are separate questions. I would not claim a cost saving without a comparable before-and-after measurement.
Changes I would test next
I would start with context management: keep a bounded recent conversation and carry verified business facts in compact application state. Historical tool JSON should not remain in every future request just because it was useful once.
I would also test exposing a smaller set of tools for each stage. A greeting should not need every inventory, pricing, approval and order schema. Any routing change would need regression tests so that lower token usage did not remove a necessary check.
Where validated state already determines the next step, Python could coordinate it without asking the model to decide again. The LLM would remain useful for interpreting language and resolving ambiguity.
These are proposed improvements, not results from this experiment. I would compare input tokens, call counts, latency and workflow correctness on the same scenarios before deciding whether they helped.
What I took from the investigation
I went into the investigation expecting conversation history to be the main problem. The fresh "Hello" baseline showed that we also needed to examine what the application sent before history had grown.
The work reinforced how closely agent behaviour depends on ordinary software design. Tool interfaces, transaction state and orchestration determine how often a model is called and what it sees. Measuring at the gateway made those effects visible.
LionCity AI SalesOps was a team-built hackathon prototype, not a production deployment. My contribution during the later integration and hardening phase included multi-turn debugging, regression testing, transaction reliability, deployment workflow and this LLM usage investigation.
Measurement note: figures are retained from the supplied gateway-diagnostics write-up. Their arithmetic has been checked for this editorial revision; the underlying diagnostic logs were not supplied for independent verification.
Related case study: A WhatsApp sales agent with business controls.