# One AI chat assistant service for several brand apps

> A queue-backed chat service that answered users inside several subscription apps: it turned bursts of messages into one reply, delivered replies in parts, and was audited after the fact by an LLM judge and SQL analytics.

Source: https://simasrazinskas.com/work/in-app-ai-assistant  
Published: 2026-10-07  
Reviewed: 2026-10-07

## At a glance

- **Problem**: A consumer subscription business selling under several brands wanted a chat assistant inside its apps, configured per brand, without each app team building its own.
- **Constraint**: Users type in bursts and keep typing while a reply is on its way, and every brand needed its own knowledge, support contact and API credentials on shared infrastructure.
- **What I built**: An HTTP service that turns bursts into one model call through a Redis batch window, stores reply parts as scheduled rows, delivers them by webhook, and isolates brands by API key. An n8n pipeline splits conversations into threads and has an LLM judge flag rule breaches.
- **Result**: A no-code prototype became a production service in two weeks and then served several brands from one deployment. Replies breaking a hard rule reached the team's chat channel on a 4-hour cycle.
- **Stack**: TypeScript · Fastify · Postgres / Drizzle · Redis / BullMQ · Vercel AI SDK · n8n · Next.js · Docker

## Overview

The business sells several subscription apps under separate brands.
It wanted a chat assistant inside those apps that users could message at any time, with answers grounded in what each brand offers.

The first version, in May 2025, was an n8n workflow: an agent node, a model call and Postgres-backed chat memory.
Within a week I rewrote it as a service, first in Node.js and then in TypeScript with Fastify, Drizzle on Postgres, and BullMQ on Redis.
The TypeScript version went to production in early June 2025.

The app calls a small REST API: create a chat, send a message, list messages.
Replies are generated by OpenAI and Anthropic models through the Vercel AI SDK, selectable per chat.
The model must return a fixed shape, `{ messages: string[] }`, so one reply arrives as a few short texts instead of one long block.
Each call gets up to 5 attempts with exponential backoff and jitter, capped at 30 seconds, on network errors, rate limits, server errors and empty or malformed output.
Other errors fail fast.

Alongside the service I built the internal dashboard: conversations, brand projects, engagement charts and an error-tracking view of everything the quality judge flagged.

## Bursts, batch windows and interruptions

People rarely send one complete message.
They send several short ones in a row, and a naive service answers each one, producing overlapping and duplicate replies.

The first fix batched messages in process memory.
Within days the batch moved into Redis, where it survives restarts.
Each chat now has a batch window: the first message creates a delayed BullMQ job and a window key with a time-to-live.
Later messages inside the window join the same job, which is re-queued with the remaining delay.
When the delay expires, a worker sends the whole burst plus the recent history to the model in one call.

The harder case was a user who writes again while the previous reply is still being delivered.
Reply parts go out one at a time, so the user's new message can arrive between parts two and three.
Without special handling, the rest of the old reply still arrived after the user's new message, and rapid messages still produced duplicate replies.

The fix treats a new message as an interruption.
Before batching, the service deletes the chat's reply parts that are not yet visible, removes their queued webhook jobs, removes any pending batch job, and clears the batch window.
The new burst then gets a 1-second window, so the reply to the latest message comes quickly.
A short Redis lock per chat guards the window bookkeeping against two messages racing through it.
The commit that shipped this closed the duplicate-reply bug.

## Delivery as data

A reply of three parts is three rows in the `messages` table, each with a `scheduled_for` timestamp spaced by its length plus a short gap.
Every read query filters out rows whose time has not come.
The worker writes all parts first, then queues one webhook per part to the app.

This made the database the source of truth for delivery.
If the app's webhook endpoint is down, the job still succeeds, because the messages are saved, and the app sees them the next time it lists the chat.
When callbacks fail, the worker tries to post a status event to the app and still marks the job done.
It also made interruption cheap: retracting an unsent part is a `DELETE` of rows with a future timestamp, with nothing to recall from the client.

The model call and the delivery run on separate queues with separate retry policies, so retrying a flaky callback never re-runs a model call.

## One service, several brands

In July 2025 the service became multi-tenant.
Each brand is a project with its own API keys, support email, product description, subscription details and knowledge text.
The bearer key on each request resolves the project, and chat and message queries are filtered by it.

The system prompt is assembled per call from four layers.
The base prompt comes with the chat.
A shared rules template holds placeholders for the project's support email, subscription details and knowledge.
Per-user onboarding answers, passed as JSON when the chat is created, are flattened into readable `Key: value` lines.
The current date is added last.
Adding a brand became a database row and a key, with no code change.

## A judge that reads every conversation

Model output needed checking after release, so I built an offline quality loop in n8n.

The unit of review is a thread: a burst of activity inside a long-running chat.
Threads are derived from message timestamps with one window-function query, the classic gaps-and-islands pattern.
A message that follows the previous one by an hour or more starts a new thread.

```sql
WITH ordered AS (
  SELECT chat_id, message_type, created_at,
         lag(created_at) OVER (PARTITION BY chat_id ORDER BY created_at) AS prev_ts
  FROM messages
),
marked AS (
  SELECT *, CASE WHEN prev_ts IS NULL
                   OR created_at - prev_ts >= interval '1 hour'
                 THEN 1 ELSE 0 END AS is_new_thread
  FROM ordered
)
SELECT chat_id,
       sum(is_new_thread) OVER (PARTITION BY chat_id ORDER BY created_at
                                ROWS UNBOUNDED PRECEDING) AS thread_no,
       created_at
FROM marked;
```

*Simplified: the production query then aggregates start, end, duration and message counts per thread.*

The same thread table drives the dashboard's engagement charts: activity by hour and weekday, thread duration distributions and retention.

Every 4 hours a workflow picks unreviewed threads from the last 3 days.
Each one is rendered as escaped XML and sent to an LLM judge with a 13-field JSON schema.
The schema covers topics, sentiment, rule warnings, rule errors, jailbreak attempts, out-of-scope requests, requests that belong with customer support, and replies that came back as refusals or empty text.
Minor slips are warnings and breaches of a hard rule are errors.
System messages are excluded from judgment.
Threads with errors are posted to the team's chat channel with a link straight to the conversation in the admin view.

The shared template kept changing while the judge ran.
Later rules forbid sharing any phone number and replace product names with criteria the user can check.
Language handling changed in the same period.
An English-only rule from September 2025 was replaced in early 2026 by replying in the user's language.

## Limitations

- In the TypeScript version the API, the AI worker and the webhook worker share one Node.js process, so a crash takes all three down. A later Go rewrite split them into separate services.
- Each chat uses one model provider, with no failover to the other when it is down.
- Webhook jobs get 3 attempts and are then discarded; there is no dead-letter queue, so a lost callback relies on the app polling.
- The worker sends up to the last 300 messages on every call, with no summarization, so long chats get slower and costlier.
- Live moderation is keyword and pattern matching that only lengthens the batch delay; the real safety rules live in the prompt.
- The per-chat lock is advisory: a second request waits 100 ms and proceeds, so a heavy burst can still race the window.
- Quality checks run after replies reach users. There was no golden test set, so a prompt change was only judged once real conversations had happened.

## Related case studies

- [A supervised AI support agent grounded in real customer data](https://simasrazinskas.com/work/ai-customer-support-agent): Support staff approve ready-made replies instead of writing every answer from scratch.
- [One governed platform for a company's internal AI automation](https://simasrazinskas.com/work/internal-ai-platform): New internal AI tools launch on one shared, access-controlled platform, not from scratch.

## Start with one process

A one-hour call costs €80. Afterwards you get a written plan, whether or not we work together.

[Book a call](https://simasrazinskas.com/book-me) · [simas@simasrazinskas.com](mailto:simas@simasrazinskas.com)
