Chat ui with multi agent and history - langraph

Hi Team

Can anyone help or add any pointers

I want to create chat ui which should call python API which should be in langraph nodes and states

Once the endpoint is called it should analyse which tool to call like , can the user prompt be answered just by calling and agent which has informed about faqs and rules .

If it cannot answer then it should call another agent or function which can answer by retrieving user info and combine with faqs and rules and return answer

Also chat will remember past user history for last couple logins and also chat remembers past 7 messages as history

Can you help if this can be done I started to use gemini embedding to create embedding and gemma for chat ui

Any reference or sample code I can refer or anyone working can you help

Looking to send info from your agent chat back to the UI? I’d definitely recommend using webhooks for this. I’ve actually implemented this exact approach in my project (https://aurocreator.com/)—happy to share more details or discuss the implementation if you’re interested!

Hmm… maybe LangGraph’s conditional workflow could be useful here?


Yes — I think the architecture you described is quite feasible in LangGraph.

I would probably start a little simpler than a full multi-agent setup, while keeping exactly the same product goal. Your flow already has a fairly clear decision boundary:

Chat UI
  -> Python API
  -> LangGraph
       -> retrieve/search FAQ + rules
       -> is that enough to answer?
            |-- yes -> answer from FAQ/rules
            `-- no
                 -> does this request need user-specific data?
                      |-- yes -> fetch authorized user data
                      |          -> combine with FAQ/rules
                      |          -> answer
                      `-- no  -> clarify / fallback

That can be implemented as one explicit LangGraph workflow first. If one branch later becomes much larger — for example it gets its own prompt, many tools, separate permissions, independent context, or genuinely independent work — that branch can become its own subgraph or specialist agent later without changing the overall design.

The current LangChain/LangGraph docs are useful here because they treat custom workflows and routers as first-class patterns. A complex application does not automatically need several autonomous agents; when the categories are fairly explicit, a conditional graph is usually easier to inspect, test, and debug.

I would separate five things from the beginning:

  1. FAQ / rules retrieval — semantic retrieval over policy-like text.
  2. User-specific lookup — account/order/ticket/profile data from an authenticated DB/API/tool path.
  3. Conversation persistence — resuming the same conversation using a stable thread_id and a checkpointer.
  4. What the model sees now — recent turns, trimming, or a summary; this is different from what you persist.
  5. Cross-login / cross-conversation memory — user-scoped durable memory only if the product actually needs it.

That separation is probably more important than deciding whether you have one agent or two agents.

Why I would start with one conditional workflow

The main reason is not that multi-agent is wrong. It is that the branches you described currently differ mostly by which information source is required, not yet by completely independent goals.

A useful boundary is:

Does this branch mainly differ by the data/tool it needs?
    -> keep it as a node/tool/conditional branch for now

Does this branch have its own large prompt/context, many tools,
separate permission boundary, or separate objective?
    -> a specialist subgraph/agent may make sense

Do several independent branches need to run in parallel
and then be synthesized?
    -> router/fan-out/multi-agent orchestration may become useful

For your case, the first route could be very small:

Route = Literal[
    "FAQ_ONLY",
    "NEEDS_USER_DATA",
    "CLARIFY",
]

Or you can keep two decisions separate:

class RouteDecision(BaseModel):
    needs_user_data: bool
    faq_evidence_sufficient: bool | None

Those are not exactly the same question.

For example:

"What is your cancellation policy?"
    -> no private user data is required

"Can I cancel order #123?"
    -> current order state is user-specific
    -> general cancellation rules may still be relevant

"Can I do this under the plan I am currently on?"
    -> current plan may need a private lookup
    -> then FAQ/rules may still be needed

So I would not force “FAQ agent vs user-data agent” to be the only conceptual split. Another useful formulation is:

Which sources are required for a correct answer?

That makes the workflow easier to evolve later.

One simple first version is FAQ-first:

question
 -> FAQ/rules retrieval
 -> enough evidence?
      |-- yes -> answer
      `-- no and user-specific facts are required
           -> authenticated user lookup
           -> combine evidence
           -> answer

If your traffic turns out to be mostly account/order/ticket requests, you can optimize later with intent-first routing:

question
 -> classify whether private user data is required
      |-- no  -> FAQ retrieval
      `-- yes -> user lookup + relevant FAQ/rules
 -> answer

The useful thing about keeping this explicit in LangGraph is that you can see where the route happened and what evidence existed at that point. If routing becomes unstable, you can move or split that decision without redesigning the entire application.

A possible evolution path is:

v1
FAQ node -> route -> user lookup node -> answer node

v2
FAQ node -> route -> customer-support subgraph -> answer node

v3
router -> multiple independent specialist subgraphs -> synthesis

So starting with one workflow does not block a more agentic architecture later.

Memory: “last 7 messages” and “past couple logins” are probably different requirements

This is one of the parts I would separate early, because “chat history” can mean several different things.

I would distinguish at least these three layers:

A. Resume the same conversation
   -> thread_id + checkpointer

B. Decide what the model sees on this turn
   -> recent turns / trimming / summary

C. Remember selected facts across conversations/logins
   -> user-scoped Store or application DB

The current LangChain/LangGraph docs make a similar distinction between short-term memory, long-term memory, and LangGraph persistence.

A. Resuming the same conversation

LangGraph persistence associates checkpoints with a thread. Conceptually:

config = {
    "configurable": {
        "thread_id": conversation_id,
    }
}

If the same conversation is invoked again with the same stable thread ID and a persistent checkpointer, the graph can resume the saved state.

If “past couple logins” means the user can sign out, return later, or the API can restart and the same conversation must still resume, an in-memory saver is not enough. You would want a persistent backing implementation.

B. “Last 7 messages” should probably be a model-context policy

I would not automatically interpret this as “store only seven messages.” You can preserve the complete conversation while presenting only a selected recent context to the model:

system instructions
+ optional older-conversation summary
+ selected long-term user memories
+ last N valid conversational turns
+ current user message

This is easier to change later and lets the UI keep a complete transcript while the model stays within a controlled context budget.

One practical detail: once you use tools, a raw slice such as:

messages[-7:]

can cut through a logical assistant tool-call / tool-result pair. So it is better to define “seven” semantically:

seven raw messages?
seven user/assistant turns?
seven valid recent turns including their tool results?
summary of older context + seven recent turns?

Any of those can be valid. The important part is that the rule is explicit and testable.

C. Cross-login / cross-conversation memory

If “remember past user history” means that a new conversation should remember selected preferences or user facts, that is closer to long-term memory than to thread history.

For example:

user 42
  preferred_language = "English"
  preferred_answer_style = "short"

That is different from persisting the transcript of conversation abc123.

I would avoid automatically copying every old message into semantic long-term memory. Old conversations can contain temporary statements, outdated information, mistakes, or sensitive text. It is usually better to decide what kinds of facts deserve to persist.

A useful decision tree is:

What should survive a new login/new conversation?

Only the ability to reopen old chats?
    -> UI/chat archive may be enough

Resume the exact same conversation/workflow?
    -> persistent checkpointer + stable thread_id

Remember a few user preferences/facts in new threads?
    -> user-scoped Store/app DB

Remember the gist of recent discussions?
    -> explicit summary/memory-writing policy

This distinction also makes debugging much easier. If something is “forgotten,” you can ask:

  1. Was it persisted?
  2. Was it stored under the correct thread/user namespace?
  3. Was it loaded for this request?
  4. Was it included in the model-visible context?
  5. Was the model expected to use it?

Those are much more actionable questions than a single “memory did not work.”

FAQ/rules retrieval and user-specific data should probably have different interfaces

Your two information sources have different semantics, so I would keep them separate even if they eventually feed the same answer node.

A useful default split is:

FAQ / rules / policy prose
    -> semantic retrieval / RAG

exact user facts
    -> authenticated DB/API/tool lookup

private unstructured user documents
    -> user-scoped RAG, if needed

Semantic search is naturally useful for questions like:

"What is the refund window?"
"What documents are required?"
"What are the eligibility rules?"

But it would not be my first primitive for questions like:

"What plan am I currently on?"
"What is the status of order 123?"
"Has my ticket been approved?"

Those are exact/current records with an authorization requirement. An ordinary database/API call is usually easier to validate because it has a clear identity, schema, and source of truth.

A route result can also be more informative than a single free-form “can the FAQ answer?” judgment. For example:

{
  "needs_user_data": true,
  "faq_evidence_sufficient": false,
  "reason": "The policy describes eligibility, but the user's current plan is required."
}

The reason can stay internal. During development it helps distinguish several failure modes:

wrong intent route
bad FAQ retrieval
insufficient FAQ evidence
missing user record
bad answer synthesis

If “user info” means private documents rather than structured account records, embeddings can still be useful. But authorization should be applied before the model receives retrieved private text:

authenticated user / tenant
 -> determine permitted document scope
 -> retrieve within that scope
 -> rank semantic matches
 -> provide authorized chunks to the model

The OWASP RAG Security Cheat Sheet is useful background for this. The practical point is simply that semantic similarity should not decide which user’s documents are visible.

User identity: keep authorization outside the model

For user-specific data, I would make identity a server/application responsibility rather than a model decision.

Conceptually:

authenticated session/token
        |
        v
trusted user_id / tenant_id
        |
        v
LangGraph runtime context
        |
        v
user-scoped DB/API/Store

The current LangChain runtime documentation supports passing request-scoped dependencies such as a user ID through runtime context. That is a good fit for this boundary.

The graph state can contain conversational/workflow state, while the trusted identity comes from the application:

@dataclass
class RequestContext:
    user_id: str
    tenant_id: str | None = None

# populated by the authenticated application layer,
# not generated by the model

Then a user-data tool can derive its allowed scope from that trusted context:

def get_my_order(order_id: str, runtime: ToolRuntime[RequestContext]):
    user_id = runtime.context.user_id
    return db.orders.find_one({
        "user_id": user_id,
        "order_id": order_id,
    })

This gives you a useful security property:

The user can ask for another user's record,
but the tool still queries only the authenticated scope.

That matters more than trying to teach the LLM “please do not access other users.”

I would also keep conversation identity and user identity separate:

thread_id / conversation_id
    -> which conversation is this?

user_id / tenant_id
    -> which data is this caller allowed to access?

One user can have several conversations, and a conversation identifier should not become an authorization credential by accident.

If you later add write actions such as “cancel order” or “update billing data,” I would separate them from read-only lookups. Read-only retrieval is a much safer first tool boundary. Actions may need confirmation, idempotency, audit logs, and stricter permission checks.

If by “chat ui” you literally mean Hugging Face Chat UI

If you just mean “a generic chat frontend,” you can ignore this section. A custom frontend can call your Python/LangGraph API directly in whatever protocol you define.

If you mean the actual huggingface/chat-ui project, there is an extra integration boundary to consider.

The current Chat UI documentation says it connects to an OpenAI-compatible API through OPENAI_BASE_URL, and chat history/users/settings/files live in MongoDB.

So one clean architecture is:

HF Chat UI
    |
    | OpenAI-compatible request/stream
    v
small API adapter
    |
    v
LangGraph

From the current Chat UI code, the backend request path also carries a ChatUI-Conversation-ID header. That suggests a convenient correlation strategy:

Chat UI conversation ID
        |
        v
LangGraph thread_id

I would treat that as an integration choice, not as a formal LangGraph requirement. It is useful because the same Chat UI conversation can map to the same graph thread, but authorization should still come from the authenticated server-side user/session.

There are two other ownership questions worth deciding if you use HF Chat UI:

Who owns conversation persistence?
Who owns tool/model routing?

Chat UI already has its own persistence and tool/router capabilities. LangGraph can also own persistence, routing, and tools. Using both is possible, but it is easier to reason about if each layer has a distinct responsibility.

For example:

Chat UI
    -> frontend transcript / conversation list

LangGraph checkpointer
    -> graph execution / conversational workflow state

application DB / Store
    -> user/account truth and durable user memory

Using the same MongoDB infrastructure does not mean those are the same logical data model.

If you want a frontend designed specifically around LangGraph rather than an OpenAI-compatible adapter, the LangGraph Agent Chat UI is another reference worth looking at.

Gemma and Gemini embeddings

I would keep these as replaceable components behind explicit interfaces rather than letting them determine the architecture.

Gemma

The exact Gemma model/runtime matters, so I would not assume all Gemma versions have identical tool-calling behavior.

For current Gemma 4, Google documents a function-calling pattern where the model proposes a structured function call, the application parses/executes it, and the tool result is returned to the model.

That matches the architecture above quite well:

model decides what information is needed
    -> application validates the requested tool/arguments
    -> application performs the authorized operation
    -> model receives the result

In other words, the model should not become the security boundary merely because it emits a function call.

For the first router, you may not need a large model at all. A deterministic rule, a small classifier, or a structured-output call can be enough if the categories are clear.

Gemini embeddings

Using Gemini embeddings for FAQ retrieval is a reasonable direction. The current Gemini embeddings documentation is worth checking against the exact embedding model you select, because task/query-document conventions and model generations can differ.

I would version your FAQ index with at least:

source document/revision
chunk ID
embedding model
index version
active/inactive status

That makes policy updates and re-embedding much easier later.

Also, retrieval quality is not only an embedding-model question. Chunk size, source metadata, rule versioning, and how you formulate the query can matter just as much.

A concrete first implementation shape

I would keep mutable workflow state and trusted runtime context distinct.

For example:

from typing import TypedDict, Literal
from dataclasses import dataclass

class GraphState(TypedDict, total=False):
    messages: list
    faq_hits: list
    route: Literal["FAQ_ONLY", "NEEDS_USER_DATA", "CLARIFY"]
    faq_evidence_sufficient: bool
    user_context: dict
    final_answer: str

@dataclass
class RequestContext:
    user_id: str
    tenant_id: str | None = None

Then the flow can stay explicit:

START
  -> retrieve_faq
  -> decide_route
       |-- FAQ_ONLY
       |     -> answer_from_faq
       |
       |-- NEEDS_USER_DATA
       |     -> fetch_authorized_user_context
       |     -> answer_with_user_context
       |
       `-- CLARIFY
             -> clarify
  -> END

A route node can produce structured output:

class RouteDecision(BaseModel):
    route: Literal["FAQ_ONLY", "NEEDS_USER_DATA", "CLARIFY"]
    faq_evidence_sufficient: bool | None = None

The user-data tool should use trusted runtime context rather than accepting arbitrary identity from the model:

def fetch_user_context(state, runtime):
    user_id = runtime.context.user_id

    # Read only the fields needed for this question.
    record = lookup_for_authenticated_user(user_id, state["messages"][-1])
    return {"user_context": record}

Then invoke the graph with two independent identifiers:

config = {
    "configurable": {
        "thread_id": conversation_id,
    }
}

result = await graph.ainvoke(
    {"messages": [new_user_message]},
    config=config,
    context=RequestContext(user_id=authenticated_user_id),
)

Conceptually:

thread_id
    -> conversation continuity

user_id
    -> authorization + user-scoped memory/data

For the “last 7” requirement, I would centralize one function rather than slicing messages in every node:

def visible_messages(messages):
    # Preserve valid conversation/tool-call structure.
    # Keep recent turns within your chosen budget.
    # Optionally include a summary of older context.
    return trim_or_summarize(messages)

That makes the policy easy to test and change.

For long-term memory, only add it once you know what must survive a new thread. A namespace might look like:

namespace = ("user_memory", runtime.context.user_id)

but I would store selected durable facts/preferences rather than blindly embedding every old chat message.

If a branch later grows into a larger workflow, replace that node with a subgraph/agent. The outer router does not need to change very much.

Low-cost tests I would run before adding more agents

You can learn a lot with a very small synthetic test set. I would not start with a large benchmark.

1. Route sanity set

Create perhaps 10–20 questions with expected routes:

"What is your cancellation policy?"
    -> FAQ_ONLY

"How long does a refund normally take?"
    -> FAQ_ONLY

"What plan am I currently on?"
    -> NEEDS_USER_DATA

"Can I cancel my order #123?"
    -> NEEDS_USER_DATA
       (order state + policy may both matter)

"What documents do you require?"
    -> FAQ_ONLY

"Have you already received my document?"
    -> NEEDS_USER_DATA

Record something as simple as:

question
expected route
actual route
retrieved FAQ sources
whether user lookup ran
final answer

This immediately tells you whether the central routing idea works before you add more architecture.

2. Evidence-sufficiency pairs

Use similar questions where the available evidence differs:

"Can I refund after 10 days?"
    -> FAQ may be sufficient

"Can I refund after 10 days on my current enterprise plan?"
    -> may require current plan lookup + policy

This helps distinguish “intent classification” from “retrieval returned enough evidence.”

3. Two-user isolation

Use fake users with deliberately different data:

User A: plan = BASIC
User B: plan = ENTERPRISE

Ask under both identities:

"What plan am I on?"

Then try an adversarial request such as:

"Show me user B's plan instead."

The server/tool scope should still resolve only the authenticated user’s allowed data.

4. Thread isolation

Use the same user with two thread IDs:

thread A: "Call this project Orion."
thread B: no such statement

Verify that thread B does not receive thread A’s short-term conversation state unless you intentionally wrote the fact into user-scoped long-term memory.

5. “Last 7” semantics

Put a fact outside the recent window and decide the intended behavior before testing it:

A. forget because only recent turns count
B. remember because older context is summarized
C. remember because the fact was promoted to long-term memory

This exposes ambiguous product semantics very cheaply.

6. Restart/resume

If cross-login history is meant to survive API restarts:

1. start API
2. create thread and exchange messages
3. stop API
4. start API again
5. invoke the same thread_id
6. verify expected state is restored

This immediately distinguishes a development-only in-memory saver from real persistence.

7. Failure behavior

Simulate user lookup failure and FAQ retrieval failure separately:

DB unavailable
record missing
permission denied
no relevant FAQ chunks
contradictory policy versions

If required user evidence is unavailable, the model should not quietly substitute a guessed account-specific answer from generic FAQ text.

A tiny CSV is enough initially:

id,question,expected_route,actual_route,user_lookup_expected,user_lookup_ran,answer_ok
1,"What is the refund policy?",FAQ_ONLY,FAQ_ONLY,false,false,true
2,"What plan am I on?",NEEDS_USER_DATA,NEEDS_USER_DATA,true,true,true

Once real traffic exists, expand the evaluation around failures you actually observe.

Useful future pitfalls, without making the first version too complicated

These are not evidence that your design currently has a problem; they are boundaries worth keeping visible.

Private state and public streaming are different concerns

If raw user/account data enters graph state, do not automatically serialize the whole state to the frontend. The graph may know more than the UI should receive.

A good rule is:

internal workflow state
    != public response payload

Return a small public projection such as the final answer/status rather than the full state object.

Checkpoints are not necessarily business truth

Conversation/workflow state and authoritative business state are different.

LangGraph state/checkpoint
    -> conversational/workflow state

account/order/ticket DB/API
    -> authoritative current business state

If an order changes from pending to shipped, re-read according to your freshness requirements rather than trusting an old remembered copy forever.

Durable memory can become stale

A remembered preference may be appropriate long-term memory; a current plan/order status usually has an authoritative source and may need re-fetching.

A useful taxonomy is:

stable preference
    -> candidate for user memory

mutable business state
    -> re-fetch from source of truth

conversation-only temporary fact
    -> thread context / summary

Multiple routing/tool layers can duplicate responsibility

If HF Chat UI performs its own model/tool routing and LangGraph also performs routing/tool execution, define which layer owns what. Otherwise the system can become difficult to reason about even if each component works individually.

Embedding/model changes are migrations worth testing

If you change an embedding model, do not assume an old vector index remains compatible. If you change the model used for routing, rerun the small fixed route test set because structured-output behavior can shift even when the new model is generally stronger.

Private RAG requires retrieval-time scope

If private documents eventually share one vector backend, carry tenant/user/ACL metadata and filter during retrieval. Do not retrieve globally and ask the LLM to ignore unauthorized chunks.

Memory is data, not system authority

Long-term memory and retrieved documents may contain stale or user-controlled text. They should not be allowed to redefine tool permissions or application policy. The OWASP AI Agent Security Cheat Sheet is a useful background reference for memory isolation and least-privilege tool access.

None of these require building a large security platform before the prototype. Keeping the boundaries explicit now simply gives you a sensible place to add controls later.

A code example that looks fairly close to your use case

One community project that is conceptually close is:

production-ai-customer-support-langchain

It combines several ideas similar to what you described:

LangGraph StateGraph
conditional routing
policy/FAQ RAG
SQLite-backed tools
customer profile/context
Gemini-related components
conversation memory

It is useful for seeing how a customer-support graph can connect routing, RAG, and database-backed tools in one project.

I would still treat it as a community code reference, not a production authority. In particular, its current persistence/memory setup should not be assumed to solve your cross-login requirements automatically, and your authentication/authorization requirements may differ.

So I would use it mainly for implementation ideas such as:

StateGraph structure
conditional edges
RAG node
DB/tool node
answer synthesis

while using current LangGraph/LangChain documentation as the authority for persistence/runtime contracts.

A compact decision tree

If I were turning your description into implementation choices without blocking on more questions, I would use something like this:

Incoming request
|
|-- Does the answer obviously require current private user/account data?
|       |
|       |-- yes
|       |    -> authenticated structured DB/API lookup
|       |    -> retrieve relevant FAQ/rules if policy is also needed
|       |    -> answer from both sources
|       |
|       `-- no / unclear
|            -> retrieve FAQ/rules
|            -> is the retrieved evidence enough?
|                 |
|                 |-- yes -> FAQ-only answer
|                 |
|                 `-- no
|                      |-- private user fact required?
|                      |       -> authenticated lookup
|                      |
|                      `-- information genuinely missing?
|                              -> clarify/fallback
|
|-- What kind of history is needed?
|       |
|       |-- same conversation -> thread_id + persistent checkpointer
|       |-- current model context -> trim/recent turns/summary
|       `-- cross-conversation facts -> user-scoped Store/app DB
|
|-- What UI?
|       |
|       |-- generic frontend -> call LangGraph API directly
|       `-- HF Chat UI -> OpenAI-compatible adapter + conversation/thread mapping
|
`-- Has one branch become independently complex?
        |
        |-- no -> keep it as a node/tool/branch
        `-- yes -> promote that branch to a subgraph/specialist agent

That lets you start now without requiring every product decision to be finalized first.

Reference links I would keep nearby

LangGraph / LangChain workflow structure

Persistence and memory

UI

Models / embeddings

Private data / RAG boundaries

Close community example

So my suggested first version would be:

1. One explicit LangGraph conditional workflow.
2. FAQ/rules as semantic retrieval.
3. Private structured user data as authenticated DB/API tools.
4. Stable thread_id + persistent checkpointer for conversation continuity.
5. Treat "last 7" as a model-context policy, not automatically a storage policy.
6. Add user-scoped long-term memory only for facts that really should cross threads.
7. If you literally use HF Chat UI, add a small OpenAI-compatible adapter and decide
   how Chat UI conversation IDs map to LangGraph thread IDs.
8. Promote a branch into its own agent/subgraph only when it develops genuinely
   independent complexity.

A tiny route test set plus a two-user isolation test would probably tell you more at this stage than adding another agent immediately.

Thank you @John6666 I will try this.

This is good to start.