SeaweedFS: A Distributed File System and S3 Store Built for Billions of Files

SeaweedFS: A Distributed File System and S3 Store Built for Billions of Files

SeaweedFS is a distributed storage system designed around one problem most file systems handle badly: very large numbers of small files. It can serve them from disk in a single seek, and it exposes the result through an S3 API, a POSIX-like mount, WebDAV, and SFTP.

Architecture

ComponentRole
MasterTracks volumes and assigns file IDs. It is not in the read path
Volume serversStore file data in large append-only volume files, with an in-memory index
FilerAdds directories and file names on top, with pluggable metadata stores

Because volume servers pack many small files into big volume files and index them in memory, reading a file is one disk seek, and per-file metadata overhead is around 40 bytes. That is why SeaweedFS holds up where storing millions of thumbnails or log chunks as individual files would crush an ordinary filesystem.

Features

  • S3-compatible object storage
  • FUSE mount on Linux, macOS, and Windows
  • WebDAV and SFTP
  • Replication for hot data and erasure coding for warm data
  • Tiering to cloud storage
  • Apache Iceberg table support with a built-in REST catalog

Deployment

SeaweedFS ships as a single weed binary that can run every component, which makes a one-machine test easy:

weed server -dir=/srv/seaweedfs -s3

Production clusters run masters, volume servers, and filers separately, using Docker Compose, Kubernetes with Helm charts, or bare metal.

When to use it

SeaweedFS shines for large deployments and small-file-heavy workloads: media thumbnails, backups with many chunks, logs, and data lakes. For a single homelab box that just needs an S3 endpoint, Garage, RustFS, or MinIO are simpler to operate.

License

Apache License 2.0.