The n8n AI starter kit is an open-source Docker Compose template that runs n8n, Ollama, Qdrant and PostgreSQL together for local AI workflows. It suits learning, prototyping and keeping data on your own machine. Its README says it is not fully optimised for production; for the strongest models or simple hosting, plain n8n with API models is often the better fit.
The n8n AI starter kit is an open-source Docker Compose template from n8n that runs n8n, Ollama, Qdrant and PostgreSQL together, giving you a local environment for building AI workflows without sending data to outside model APIs. It is designed for learning and proof-of-concept work; its README says it is not fully optimised for production. Whether you need it depends on whether keeping models local matters more to you than model quality and simple setup.
Checked October 2026 against the self-hosted-ai-starter-kit repository and its docker-compose.yml, and n8n's docs page on deploying with the AI starter kit.
The kit is one of many ways to run n8n. For templates, integrations and other hosting options, see our n8n hub with 100 workflow templates.
n8n AI starter kit: what is in the box
| Component | What it is | Its job in the kit |
|---|---|---|
| n8n | Workflow automation and AI agent builder | Where you build the workflows |
| Ollama | Runs open language models locally | The model your workflows call |
| Qdrant | Vector database | Stores document embeddings for search by meaning |
| PostgreSQL | Relational database | n8n's own database for workflows and data |
| Docker Compose | Runs all four as containers | One command to start the stack |
- 01n8n workflow
Trigger, steps, AI agent
- 02Qdrant
Finds relevant passages
- 03Ollama
Local model answers
- 04PostgreSQL
Stores n8n data
- 05Your machine
Nothing leaves it
In the compose file, n8n is configured to use PostgreSQL as its database, and an Ollama host variable points n8n at the Ollama container by default. Each service gets its own Docker volume, so workflows, model downloads and vector data persist between restarts. The README says the kit includes a starter workflow, and that on first start Ollama downloads Llama 3.2, so the first run takes a while.
How to run the n8n self-hosted AI starter kit
You need Git and Docker with Docker Compose. Clone the repository, copy the example environment file, set your own secrets and database credentials, then start the stack with the profile that matches your hardware.
| Hardware | Command | Note |
|---|---|---|
| Nvidia GPU | docker compose --profile gpu-nvidia up | Follow Ollama's Docker GPU instructions first |
| AMD GPU on Linux | docker compose --profile gpu-amd up | Linux only |
| CPU only | docker compose --profile cpu up | Works anywhere, slower models |
| Apple Silicon Mac | CPU profile, or Ollama installed natively | Set OLLAMA_HOST to host.docker.internal:11434 for native Ollama |
- 1Clone the repository
n8n-io/self-hosted-ai-starter-kit on GitHub.
- 2Create your .env file
Set database credentials, the n8n encryption key and the JWT secret. Never commit it.
- 3Start with your profile
gpu-nvidia, gpu-amd or cpu.
- 4Wait for the model download
Ollama pulls Llama 3.2 on first start.
- 5Open n8n and the starter workflow
Use the Ollama node for the model and Qdrant as the vector store to stay local.
Apple Silicon Macs are the special case. The README explains that Docker cannot use the Mac GPU, so you either run the kit on CPU or install Ollama directly on the Mac for faster inference and point the kit at it by setting OLLAMA_HOST to host.docker.internal:11434. For general Docker setup, including reverse proxies and backups, see our n8n Docker Compose setup guide.
n8n Ollama starter kit: what local models change
The defining feature of the kit is Ollama. Running models locally changes three things compared with calling a hosted API.
- Prompts and documents stay on your machine
- No per-call API bill
- Speed depends on your hardware
- Open models, usually smaller
- You maintain the stack
- Data is sent to the provider
- Pay per token or per call
- Fast regardless of your hardware
- Access to the largest frontier models
- Provider maintains the models
Neither column is better for every job. Local models make sense for private documents, offline work and experimenting without a bill. Hosted models make sense when quality matters most or your hardware is limited. n8n works with both, so you can switch a workflow's model node later without rebuilding the rest. Our guide to n8n AI integrations with OpenAI covers the hosted side.
What you can build with it
The kit is set up for retrieval-augmented generation and AI agents. Typical first projects include:
- Private document Q&A: embed your PDFs into Qdrant, then ask questions answered by the local model with relevant passages.
- Internal chat assistant: an n8n chat trigger with an AI agent that searches your knowledge base.
- Classification and extraction: sort incoming emails or files and pull out fields without sending them to a third party.
- Agent experiments: compare agents and chains on the same task, which the README lists as one of the things the kit demonstrates.
For the concepts behind agents and multi-step AI flows, see our guide to agentic AI workflows.
Hardware and resource planning
Running four services and a language model on one machine is heavier than running n8n alone. The model is the biggest consumer: Ollama loads model weights into memory, and larger models need more RAM or GPU memory and run more slowly on modest hardware. Qdrant and PostgreSQL add their own memory and disk use as your document collection grows.
A sensible approach is to start with the default model the kit downloads, measure how long a typical workflow takes on your machine, and only then decide whether you need a bigger model, a GPU, or a hosted model for some steps. Ollama's own documentation lists the models it supports and their sizes, which is the place to check before pulling something large.
- Disk: model downloads, vector data and database volumes all persist in Docker volumes, so leave room for growth.
- Memory: the model, the vector store and n8n share the same RAM on a single machine.
- Speed: CPU-only inference works, but expect slower responses than on a GPU or a hosted API.
Mixing local and hosted models
You do not have to choose one approach for everything. Because each AI step in n8n is a node, a single workflow can use a local Ollama model for steps that touch private data, such as classifying internal documents, and a hosted model for steps where quality matters most and the data is not sensitive. Qdrant can stay local in both cases. This hybrid pattern keeps sensitive text on your machine while still giving you access to stronger models where you need them.
Is it production-ready?
The README is direct: the kit is not fully optimised for production environments and combines components that work well together for proof-of-concept projects. Before relying on a stack like this for real users or business data, plan for:
- Strong, unique secrets in .env, kept out of version control
- Pinned image versions instead of latest
- Backups of the n8n, PostgreSQL and Qdrant volumes
- HTTPS and authentication in front of n8n
- Resource monitoring for CPU, RAM and disk
- Error handling and alerts on key workflows
- A plan for model updates and testing
The compose file uses latest image tags, which is convenient for a demo but means an update can change behaviour unexpectedly. Pinning versions is one of the first changes to make. For keeping workflows reliable, see our guide to n8n error handling, and for hosting costs, our breakdown of n8n self-hosted costs.
Do you need the starter kit?
- Use it if you want to learn local AI with n8n, need documents to stay on your machine, or want to prototype RAG and agents without API costs.
- Skip it if you only need n8n with hosted models, have limited RAM or no GPU and need fast responses, or want a production setup today. Plain n8n, self-hosted or on n8n Cloud, with an API model is simpler.
If you want to turn automations like these into products or services, our AI SaaS program covers building and selling AI tools step by step.
n8n AI starter kit: FAQ
What is the n8n AI starter kit?
The Self-hosted AI Starter Kit is an open-source Docker Compose template from n8n that sets up a local AI development environment. It bundles n8n with Ollama for running language models locally, Qdrant as a vector store, and PostgreSQL for data. n8n describes it as a way to start building self-hosted AI workflows quickly. Checked October 2026.
Is the n8n AI starter kit ready for production?
Not as it ships. Its README says it is not fully optimised for production environments and is aimed at proof-of-concept projects. Treat it as a learning and prototyping stack. For production, you would harden security, set proper secrets, plan backups, monitor resources, and usually use a separate production n8n setup.
Do I need a GPU to run the n8n starter kit?
No. The kit has Docker Compose profiles for Nvidia GPUs, AMD GPUs on Linux, and CPU only. Without a GPU, models run on the CPU, which is slower. On Apple Silicon Macs, the README notes you cannot expose the GPU to Docker, so it suggests running on CPU or running Ollama natively on the Mac and connecting to it.
What does Ollama do in the n8n starter kit?
Ollama runs open language models on your own machine, so prompts and documents do not leave it. The kit's README notes that on first start it downloads Llama 3.2, and the included workflow waits until that finishes. In n8n, you use the Ollama nodes as your language model to keep everything local.
Should I use the starter kit or plain n8n with API models?
Use the starter kit if you need data to stay on your machine, want to learn local AI, or want to avoid per-call API costs during experiments. Use plain n8n with hosted API models if you need the strongest models, have modest hardware, or want the simplest setup. Many teams prototype locally and run production on hosted models.
What is Qdrant used for in n8n?
Qdrant is a vector database. In the starter kit it stores embeddings of your documents so an AI workflow can search by meaning and pass relevant passages to the model, which is the core of retrieval-augmented generation. The README says to use Qdrant as your vector store to keep everything local.
From local prototype to a product people use.
The AI SaaS program, included in All Access, covers automation, AI agents and building sellable AI tools, with the other three programs, live coaching and the private community in one subscription.
Free n8n help
Join the free Discord to share workflows, ask setup questions and see what others are building.