— Independent data platform consultancy
Sensor to security, every number with a venue and a date.
Pipelines, platforms, AI agents, and the applications on top. Built on Databricks and Snowflake, for the teams that run them. You can check the work before you hire us.
01 Source & stream
Late data has rules.
Where the pipeline starts: a real sensor on a real board. The MAX30102 below is the soldered-for-real build from the streaming work, wired over I²C, its signal feeding the Structured Streaming miniature beneath it — tumbling windows keyed on event time, a watermark deciding what merges and what drops.
the hardware · a MAX30102 on a Raspberry Pi — where the stream begins
from the real-time IoT build · the published write-up
engage at this stage · Pipeline builds, watermark semantics, Auto Loader design · see the practice →
02 Platform & benchmarking
We run the benchmark. You keep the harness.
The stream above just committed its rows; keeping them fast and affordable is the next problem. Compute selection, cost modeling, platform bake-offs: we design the workload, run it under controlled conditions, and hand over an instrument your team can re-run without us. The published record is the credential — including the findings that cut against the obvious answer.
- 4,750+ TPC-DS query executions in the published record
- 55.8B rows in the largest test environment
- 2 × 6 platforms × compute configurations benchmarked in print
- Yours the harness ships with the findings — re-run it after we leave
engage at this stage · Benchmarks, compute-plane selection, cost modeling · see the practice →
supporting evidence 01 · a published finding — snowflake gen2 vs gen1, tpc-ds 1tb
On TPC-DS at 1TB with small warehouses, a Gen2 warehouse finished in 0.64 minutes where Gen1 took 38.24, while consuming 9.66 credits against Gen1's 10.58.
TPC-DS 1TB, cold start, small warehouses. Same size, nearly the same credits — but a 60× runtime gap driven by memory spilling. At large sizes, with no memory pressure, the result inverts and Gen1 wins the majority of queries.
source · Snowflake Gen1 vs. Gen2 vs. Snowpark-optimized warehouses · February 5, 2026
View the data as a table
| Measure | Warehouse | Value |
|---|---|---|
| Runtime | Gen2 small | 0.64 |
| Runtime | Gen1 small | 38.24 |
| Credits consumed | Gen2 small | 9.66 |
| Credits consumed | Gen1 small | 10.58 |
supporting evidence 02 · the instrument — reproducibility, live
Measured and paid for — now people need to work with it. ↓ 03 · applications & bi
03 Applications & BI
Clean, cutting-edge apps — like the ones running on this page.
A measured platform still needs its human layer. The claim there is craft, and the proof is local: every interactive on this site is built the way we build client applications — hand-written, token-themed, honest to the pixel. Below, the two halves of the human layer as exhibits: reading at any scale, and writing with an audit trail.
- 8 live canvases running on this page — demos and paintings
- 0 frameworks, libraries, or external requests
- 81 KB of JavaScript — all eight, total
- <2ms per animation frame, measured
- Faithful in both themes and under reduced motion
exhibit 01 · reading — the resampler: billions of points, live
the technique from the billion-point write-up and the Mercedes engagement · move across it
exhibit 02 · writing — governed write-back: every edit an auditable transaction
from the supply-planning engagement · read the case
engage at this stage · Warehouse-backed applications and BI-layer builds · see the practice →
Humans read and write; agents need something more — context. ↓ 04 · ai agents
04 AI agents
Context is a graph, not a keyword.
Humans get dashboards; agents need context. The argument from the knowledge-graph study — and the reason we built lakecode, our own assistant with compiled, persistent context. Click any entity and compare flat retrieval with a bounded graph walk.
The platform · lakecode.ai
lakecode
Our AI assistant and agent platform, and the thing the argument above is built out of. Its context is compiled into a persistent memory substrate rather than re-fetched each session, so what it knows about your systems accumulates, survives across sessions, and is governed like any other data asset. It speaks MCP, so your team keeps whatever agent front-end it already uses.
We offer it to clients who need it — deployed against your warehouse and your codebase, rather than rebuilding the same substrate from scratch inside an engagement. Where it is not the right fit, the architecture work stands on its own and we will say so.
from the knowledge-graph study · read it
engage at this stage · Agent architecture, context engineering, and lakecode deployed on your stack · see the practice →
Every stage above moves data that has to stay protected. ↓ 05 · security & tokenization
05 Security & tokenization
Protect it at the source and it stays protected.
The layer that crosses every stage: data protection that survives the workload. Our July 2026 study argues for tokenizing before the API boundary — written as Claude Fable 5 shipped with a 30-day retention window. The follow-up benchmark, tokenization vs. masking vs. deterministic encryption across five experiments, is in progress and will publish like everything else here.
engage at this stage · Tokenization architecture and scheme selection · see the practice →
That is the stack, and every stage of it was demonstrated rather than described. ↓ — · the research
— The research
Every number above has a paper behind it.
The five stages above are demonstrations, not claims — each one runs on this page, and you can watch it work before you believe it. What stands behind them is published: 13 studies and a conference talk, every figure carrying a venue and a date, alongside the client systems those figures were measured on.
— Contact
Start with the problem.
No discovery deck, no junior on the call. A paragraph about what your platform is doing, or should be doing, is enough to begin — and we will say plainly if it isn’t work we should take.
sachin@lakesideanalytics.io