// the find
graphite-project/whisper
Whisper is a file-based time-series database format for Graphite.
Whisper is the fixed-size, RRD-style time-series storage format used by Graphite — one file per metric, with configurable retention tiers that downsample older data. It's for people running or maintaining a Graphite stack, not for anyone picking a time-series DB from scratch today.
The file format is dead simple and predictable: pre-allocated size, fixed-width records, so disk usage per metric is known in advance and reads/writes are O(1) seek-based rather than requiring an index. The CLI toolset (whisper-resize, whisper-merge, whisper-fill, whisper-diff, rrd2whisper) covers real operational pain points like migrating retention schemes or reconciling two files after a split-brain, which is the kind of thing you only appreciate once you've run this in production. It's been battle-tested for over a decade inside Graphite deployments, so the edge cases in the aggregation/rollup logic are mostly shaken out.
One file per metric means tens of thousands of metrics turn into tens of thousands of small files — that's an inode and IOPS problem on spinning disk or networked storage, and it's the main reason people eventually move to Cassandra- or InfluxDB-backed Graphite setups at scale. There's no compression: sparse or flat metrics still consume their full pre-allocated size, so fixed retention plans can waste significant disk. The codebase still leans on optparse instead of argparse and the README just points to readthedocs for actual usage instead of documenting it inline, which tells you docs aren't a priority and may be stale. It's also a single-writer-per-file design with no built-in replication or sharding — durability and scale-out are entirely someone else's problem (Carbon's).