Skip to content
Donné Alphonse.

Donné Alphonse SOLOFONDRAIBE

Python Developer · Data Architect · AI Engineer

I design reliable data platforms and AI systems, from ingestion to LLM serving.

years of experience
4+
app/api.py
from typing import Annotated from fastapi import Depends, FastAPIfrom fastapi.responses import StreamingResponsefrom pydantic import BaseModel from app.auth import Tenant, current_tenantfrom app.llm import stream_tokensfrom app.retrieval import build_prompt, search app = FastAPI(title="RAG API")CurrentTenant = Annotated[Tenant, Depends(current_tenant)]  class Question(BaseModel):    text: str  @app.post("/v1/ask")async def ask(q: Question, tenant: CurrentTenant):    hits = await search(q.text, tenant_id=tenant.id, k=5)    prompt = build_prompt(q.text, hits)     async def events():        async for token in stream_tokens(prompt):            yield f"data: {token}\n\n"     return StreamingResponse(        events(), media_type="text/event-stream"    )
Generic example: tenant-filtered retrieval, then the LLM answer streamed over SSE.
01

Positioning

Areas of expertise

A Python, data and AI profile, focused on designing reliable platforms rather than isolated demos.

  • Python Engineering
  • Data Architecture
  • Data Engineering · ETL / ELT
  • Data pipelines
  • LLM / Generative AI
  • RAG
  • AI Infrastructure
  • Backend Engineering
  • API / Microservices
  • PostgreSQL
  • Vector Databases
  • Docker
  • Security
03

Data & AI Architecture

How I design systems

Reference architectures, from ingestion to LLM serving — on-premise or via API. Hover or tab through the blocks for details, then read the decisions and their trade-offs.

Integrating an LLM: on-premise or via API

9-step flow: Client, then FastAPI API, then Tenant isolation, then RAG pipeline, then Qdrant · pgvector, then PostgreSQL, then LLM inference, then Streaming, then Monitoring · logs. Use the arrow keys to move from one step to the next.

  1. · Business application

    The business software consumes the AI engine through an API. Data stays inside the company's infrastructure.

  2. · JWT auth

    Single entry point: JWT authentication, request validation and integration endpoints.

  3. · Access control

    Every request is confined to its tenant's scope: access control and data isolation between customers.

  4. · Retrieval · prompt

    Semantic search over the business knowledge base, then assembly of the context and prompt sent to the model.

  5. · Vector index

    Index of chunked, embedded documents, queried by similarity for semantic search.

  6. · Relational DB

    Application relational data, read by the pipeline alongside vector search.

  7. · vLLM · on-premise

    Language model served locally with vLLM: no data is sent to an external cloud service.

  8. · Token by token

    The answer is sent back to the client as it is generated, which strongly reduces perceived latency.

  9. · Perf · logs · security

    Supervision of the whole chain: logs, performance and security events.

The business software consumes the AI engine through an API. Data stays inside the company's infrastructure.

Single entry point: JWT authentication, request validation and integration endpoints.

Every request is confined to its tenant's scope: access control and data isolation between customers.

Semantic search over the business knowledge base, then assembly of the context and prompt sent to the model.

Index of chunked, embedded documents, queried by similarity for semantic search.

Application relational data, read by the pipeline alongside vector search.

Language model served locally with vLLM: no data is sent to an external cloud service.

The answer is sent back to the client as it is generated, which strongly reduces perceived latency.

Supervision of the whole chain: logs, performance and security events.

Generic view of the sovereign AI engine: from the client request to the streamed answer, without any data leaving the infrastructure.

Architecture decisions

  • ADR-A1

    Local LLM inference instead of a cloud API
    Context
    Business data is sensitive and the software targets regulated environments: it must not leave the infrastructure.
    Decision
    Serve open-weight models locally with vLLM, behind the internal API.
    Trade-offs
    GPU capacity must be sized and operated in-house; choice limited to open-weight models, sometimes behind proprietary models on some tasks; model and driver upgrades are the team's responsibility.
  • ADR-A3

    Token-by-token response streaming
    Context
    A full LLM generation can take several seconds; waiting for completion badly hurts the user experience.
    Decision
    Stream tokens from inference to the client through a streamed HTTP response.
    Trade-offs
    Errors can happen mid-stream and must be surfaced cleanly to the client; reverse-proxy buffering and timeouts (Nginx) must be tuned; logging and testing are harder than with a single response.

On-premise vs API comparison

  • Data sovereignty
    Sovereign on-premise
    Data stays inside the infrastructure.
    Via API (API key)
    The prompt and its context go through the provider; retention and location must be checked contractually.
  • Compliance
    Sovereign on-premise
    Fully controlled perimeter, suited to regulated environments.
    Via API (API key)
    Depends on the provider: data processing agreement, hosting region, certifications.
  • Maintenance
    Sovereign on-premise
    Models, drivers and GPU capacity are the team's responsibility.
    Via API (API key)
    Models maintained by the provider; keys, quotas and API version changes to manage.

End-to-end data pipeline

8-step flow: Sources, then Ingestion, then Transformation, then Embedding, then Qdrant · pgvector, then PostgreSQL, then Serving API, then Logs · monitoring. Use the arrow keys to move from one step to the next.

  1. · MySQL · PgSQL · Web

    MySQL / PostgreSQL relational databases and web content.

  2. · ETL / ELT · scraping

    ETL/ELT extraction and loading from databases; scraping (SeleniumBase) for web sources.

  3. · Cleaning · Pandas

    Data cleaning, normalisation and enrichment with Pandas.

  4. · Embeddings

    Content chunking and embedding computation for semantic search.

  5. · Vector store

    Embedding index, queried by similarity.

  6. · Relational store

    Structured, cleaned data served to application queries.

  7. · FastAPI

    FastAPI API exposing search and data to applications and AI agents.

  8. · Runs · errors · latency

    Execution logs and monitoring of pipelines and the API.

MySQL / PostgreSQL relational databases and web content.

ETL/ELT extraction and loading from databases; scraping (SeleniumBase) for web sources.

Data cleaning, normalisation and enrichment with Pandas.

Content chunking and embedding computation for semantic search.

Embedding index, queried by similarity.

Structured, cleaned data served to application queries.

FastAPI API exposing search and data to applications and AI agents.

Execution logs and monitoring of pipelines and the API.

Generic reference architecture built from the components used in my engagements: from raw sources to API serving.
04

Projects

Selected projects

What each project achieves — problem, solution and role — rather than a plain list of technologies.

6 projects

  • 2026In progressFlagshipAI Infrastructure

    Sovereign AI engine

    Contribution to a self-hosted AI engine for regulated business software: RAG, local LLM inference, secured APIs and multi-tenant isolation.

    Role · AI Engineer — backend, RAG, application security and production integration

  • Agentic AI

    Business AI agents via API

    AI agents integrated into business applications and backed by an external LLM provider called with an API key, following the “Business AI via API” architecture shown above.

  • 2024CompletedAgentic AI

    JuriBot Mada Intelligent

    AI platform making the law accessible in Madagascar. Intelligent answers to legal questions in Malagasy and French.

    Role · Full-stack AI developer

  • 2025PrototypeHR & Training AI

    AI-Recruteur — Interview simulator

    Conversational agent that simulates a job interview from a job posting, generates technical and HR questions and evaluates answers.

    Role · AI developer

  • 2025CompletedAgentic AI

    AI agent for data scraping

    Agentic AI system that scrapes web content and answers questions from the collected data through semantic search.

    Role · Python AI developer

  • 2024CompletedBackend / ERP

    Odoo ERP modules & AI integration

    Custom Odoo modules automating business processes, with an API connecting the ERP to external AI models.

    Role · Python / Odoo developer

05

Skills

Technical skills

Organised by engineering domain — only technologies I have actually used in my career.

Hover or select a skill to see where it was used.

Python Engineering

APIs, backend services and business integrations in Python.

Data Engineering

Ingestion, transformation and data pipelines from source to storage.

  • Pandas

Data Architecture

Modelling and storage choices: relational, vector and cache.

AI / LLM Engineering

RAG, agents, language model integration and serving.

  • SpaCy

Platform / DevOps

Containerisation, CI/CD, deployment and supervision.

  • Linux
  • Nginx
  • GPU inference

Security

Authentication, authorisation and data isolation.

  • SSO / token exchange
06

Experience

Professional background

Roles, responsibilities and technologies — a timeline focused on Python, data and AI engineering.

  1. Responsibilities

    • Architecture of AI systems tailored to business needs
    • Design of RAG systems, knowledge bases and vectorisation pipelines
    • Secured AI APIs for application integration
    • Prompt engineering and intelligent request routing
    • Supervision (monitoring, logs, security) and performance optimisation
08

Method

From idea to production

An end-to-end system approach: from business framing to monitoring, not just a prototype.

7-step flow: Problem, then Architecture, then Development, then AI Integration, then Testing, then Deployment, then Monitoring. Use the arrow keys to move from one step to the next.

  1. Business framing & constraints

  2. System & data design

  3. APIs, services, UI

  4. LLM, RAG, agents

  5. Quality & isolation

  6. Docker, CI/CD, on-prem

  7. Performance, logs, security

09

About

Python, data & AI engineer, production-minded

Software engineer specialised in machine learning and AI, with more than 4 years of backend development experience. I build high-performance APIs, data pipelines and AI services shipped with Docker, CI/CD and a production-minded architecture.

My current work focuses on self-hosted AI engines (RAG, LLMs, multi-tenant isolation, security) integrated into regulated business software. I favour reliability, clear architecture and measurable deliverables over throwaway prototypes.

Download resume

Education

  • Master's degree in Computer Science — Information Systems Management

    E-media Madagascar

  • Bachelor's degree in Computer Science — Internet / Intranet Application Development

    EMIT

Languages

French · Technical English

10

Contact

Let's talk about your next data or AI platform

A technical question, a project, or an architecture need — write to me.