Skip to content
back to projects

// project ·

Project Aeon: Local-First Assistant Platform

Built a private local AI assistant without sending work to cloud. Stack: FastAPI, ChromaDB, Ollama, Vue.

FastAPI / ChromaDB / Ollama / RAG / Vue 3 / TypeScript

Role: Personal project · Year: 2025 · Status: shipped

tldr: A local-first RAG platform, so your documents never leave the machine. FastAPI orchestrates, ChromaDB stores and searches embeddings, Ollama runs the model. The plumbing was easy. Getting retrieval quality right with small local models was the work.

Problem

Cloud AI assistants need your documents on someone else’s server. I wanted RAG conversations over local files with everything on my own machine: no API keys, no data leaving the host.

Constraints

  • Local models are small. They can’t brute-force relevance with a huge context window the way hosted models can.
  • Context is scarce. Every retrieved chunk competes for the same limited prompt budget.
  • Models change. Swapping the LLM should be configuration, not a code change.

Decisions

  • FastAPI as the orchestrator. It takes the query, searches the vector store, assembles the prompt, and forwards generation to the local model.
  • ChromaDB for vector storage and similarity search over document embeddings.
  • Ollama runs the LLM, so model choice is a deployment concern.
  • Vue 3 with Naive UI for the chat interface. Most of the work sits behind the API.

Outcome

The chunker turned out to be the load-bearing decision. Too small and retrieval fragments related content. Too large and the context window fills with noise. The project is complete: it answered the questions about local retrieval I built it to answer.