Skip to content
Donné Alphonse.

AI Infrastructure2026

Sovereign AI engine

Contribution to a self-hosted AI engine for regulated business software: RAG, local LLM inference, secured APIs and multi-tenant isolation.

Role
AI Engineer — backend, RAG, application security and production integration
Team
Product / engineering team
Duration
Ongoing
Status
In progress

Confidential company project — deliberately generic description, no proprietary details.

Context

Contribution, as an AI Engineer, to a self-hosted AI engine integrated into business software used in regulated environments. Users need AI assistance, but sensitive data cannot be sent to external cloud services.

Constraints

  • Sovereignty: no business data leaves the infrastructure (on-premise inference and storage).
  • Multi-tenant: strict data isolation and access control between customers.
  • Security and compliance: authentication, isolation and integration into existing software.
  • Experience: streamed answers to limit perceived latency.

Architecture

9-step flow: Client, then FastAPI API, then Tenant isolation, then RAG pipeline, then Qdrant · pgvector, then PostgreSQL, then LLM inference, then Streaming, then Monitoring · logs. Use the arrow keys to move from one step to the next.

  1. · Business application

    The business software consumes the AI engine through an API. Data stays inside the company's infrastructure.

  2. · JWT auth

    Single entry point: JWT authentication, request validation and integration endpoints.

  3. · Access control

    Every request is confined to its tenant's scope: access control and data isolation between customers.

  4. · Retrieval · prompt

    Semantic search over the business knowledge base, then assembly of the context and prompt sent to the model.

  5. · Vector index

    Index of chunked, embedded documents, queried by similarity for semantic search.

  6. · Relational DB

    Application relational data, read by the pipeline alongside vector search.

  7. · vLLM · on-premise

    Language model served locally with vLLM: no data is sent to an external cloud service.

  8. · Token by token

    The answer is sent back to the client as it is generated, which strongly reduces perceived latency.

  9. · Perf · logs · security

    Supervision of the whole chain: logs, performance and security events.

The business software consumes the AI engine through an API. Data stays inside the company's infrastructure.

Single entry point: JWT authentication, request validation and integration endpoints.

Every request is confined to its tenant's scope: access control and data isolation between customers.

Semantic search over the business knowledge base, then assembly of the context and prompt sent to the model.

Index of chunked, embedded documents, queried by similarity for semantic search.

Application relational data, read by the pipeline alongside vector search.

Language model served locally with vLLM: no data is sent to an external cloud service.

The answer is sent back to the client as it is generated, which strongly reduces perceived latency.

Supervision of the whole chain: logs, performance and security events.

Generic view of the sovereign AI engine: from the client request to the streamed answer, without any data leaving the infrastructure.

Decisions & trade-offs

Architecture decisions

  • ADR-A1

    Local LLM inference instead of a cloud API

    Context
    Business data is sensitive and the software targets regulated environments: it must not leave the infrastructure.
    Decision
    Serve open-weight models locally with vLLM, behind the internal API.
    Trade-offs
    GPU capacity must be sized and operated in-house; choice limited to open-weight models, sometimes behind proprietary models on some tasks; model and driver upgrades are the team's responsibility.
  • ADR-A3

    Token-by-token response streaming

    Context
    A full LLM generation can take several seconds; waiting for completion badly hurts the user experience.
    Decision
    Stream tokens from inference to the client through a streamed HTTP response.
    Trade-offs
    Errors can happen mid-stream and must be surfaced cleanly to the client; reverse-proxy buffering and timeouts (Nginx) must be tuned; logging and testing are harder than with a single response.

Results

  • Production-oriented on-premise architecture
  • Multi-tenant isolation
  • Streamed LLM responses

Stack

  • Python
  • FastAPI
  • LLM
  • RAG
  • Qdrant
  • PostgreSQL
  • vLLM
  • Docker
  • JWT
  • Streaming