Plotline reposted this
We removed the most important column in our database. Everything got 7x faster. Plotline processes billions of events every day across several consumer apps in the US, Southeast Asia, and India. All of it lands in ClickHouse and needs to be queryable in real time - segments, cohorts, funnels, analytics. For years, every event's properties lived in a single JSON column. One blob per row. It was 80% of our entire table - and every query had to parse it from scratch, tens of millions of times, on every single request. ClickHouse's indexing, compression, data-skipping - all blind to it. Ishan from our team went deep on this. Tried five different approaches, rigorously benchmarked each one, and landed on a design we call "positional columns" - 100 real columns that ClickHouse can actually see, index, and compress, while preserving the schemaless contract our SDK depends on. The results on our benchmark set: → Data read off disk: 940 MB → 140 MB (7x less) → CPU time: 2.5s → 0.5s (5x less) → Storage: cut in half Ishan wrote up the full technical journey - every approach we tried, why compressing JSON harder wasn't the answer, and how positional columns work in production. https://lnkd.in/gJap7zxe If you're running ClickHouse with JSON or Map columns at scale, this will save you real pain. And if you've taken a different approach to schemaless event data in column stores, we'd love to hear about it.