SeaweedFS: A Distributed File System and S3 Store Built for Billions of Files
SeaweedFS is a distributed storage system designed around one problem most file systems handle badly: very large numbers of small files. It can serve them from disk in a single seek, and it exposes the result through an S3 API, a POSIX-like mount, WebDAV, and SFTP.
Architecture
| Component | Role |
|---|---|
| Master | Tracks volumes and assigns file IDs. It is not in the read path |
| Volume servers | Store file data in large append-only volume files, with an in-memory index |
| Filer | Adds directories and file names on top, with pluggable metadata stores |
Because volume servers pack many small files into big volume files and index them in memory, reading a file is one disk seek, and per-file metadata overhead is around 40 bytes. That is why SeaweedFS holds up where storing millions of thumbnails or log chunks as individual files would crush an ordinary filesystem.
Features
- S3-compatible object storage
- FUSE mount on Linux, macOS, and Windows
- WebDAV and SFTP
- Replication for hot data and erasure coding for warm data
- Tiering to cloud storage
- Apache Iceberg table support with a built-in REST catalog
Deployment
SeaweedFS ships as a single weed binary that can run every component, which makes a one-machine test easy:
weed server -dir=/srv/seaweedfs -s3
Production clusters run masters, volume servers, and filers separately, using Docker Compose, Kubernetes with Helm charts, or bare metal.
When to use it
SeaweedFS shines for large deployments and small-file-heavy workloads: media thumbnails, backups with many chunks, logs, and data lakes. For a single homelab box that just needs an S3 endpoint, Garage, RustFS, or MinIO are simpler to operate.
License
Apache License 2.0.