9-step flow: Client, then FastAPI API, then Tenant isolation, then RAG pipeline, then Qdrant · pgvector, then PostgreSQL, then LLM inference, then Streaming, then Monitoring · logs. Use the arrow keys to move from one step to the next.
The business software consumes the AI engine through an API. Data stays inside the company's infrastructure.
Single entry point: JWT authentication, request validation and integration endpoints.
Every request is confined to its tenant's scope: access control and data isolation between customers.
Semantic search over the business knowledge base, then assembly of the context and prompt sent to the model.
Index of chunked, embedded documents, queried by similarity for semantic search.
Application relational data, read by the pipeline alongside vector search.
Language model served locally with vLLM: no data is sent to an external cloud service.
The answer is sent back to the client as it is generated, which strongly reduces perceived latency.
Supervision of the whole chain: logs, performance and security events.