Tabby: A Self-Hosted Coding Assistant
Tabby provides code completion from a model you run yourself. Editor extensions talk to your server rather than to a third-party service.
Why this is the AI use case with the clearest argument
Code completion means sending your source code, including surrounding context, to whatever is answering.
For personal projects that may be acceptable. For client work under contract, proprietary systems, or anything with a confidentiality obligation, it is frequently not permitted, and many organisations have blocked hosted assistants for exactly this reason.
Tabby removes the question. The model runs on your hardware and nothing leaves.
What it does
Inline completion as you type, and a chat interface for questions about the code you have open. It can index a repository so completions reflect your codebase rather than generic patterns, which is the difference between useful and generic suggestions.
Extensions exist for VS Code, the JetBrains IDEs, Vim and Neovim.
Hardware
Code models are smaller than general chat models, so the requirements are lower than you might expect. A GPU with 8GB of VRAM handles a capable model comfortably.
CPU inference works and completion latency becomes noticeable enough to be irritating, since the value of inline completion depends on it appearing before you have typed the next line yourself.
Realistic expectations
A self-hosted model of a size you can actually run will not match the largest commercial assistants. It is good at boilerplate, repetitive patterns, and completing something you have clearly started.
The honest comparison is not against the best hosted assistant but against no assistant, which is what many people in regulated environments actually have.
Running it
A single container, with GPU passthrough if you have one. It exposes a web interface for administration and an API the extensions use.
Team deployments support multiple users with individual access tokens.
License
Tabby is released under the Apache License 2.0.