— Independent data platform consultancy

Sensor to security, every number with a venue and a date.

Pipelines, platforms, AI agents, and the applications on top. Built on Databricks and Snowflake, for the teams that run them. You can check the work before you hire us.

system starting… events · windows · warehouse · agent live · synthetic data

01 Source & stream

Late data has rules.

Where the pipeline starts: a real sensor on a real board. The MAX30102 below is the soldered-for-real build from the streaming work, wired over I²C, its signal feeding the Structured Streaming miniature beneath it — tumbling windows keyed on event time, a watermark deciding what merges and what drops.

the hardware · a MAX30102 on a Raspberry Pi — where the stream begins

from the real-time IoT build · the published write-up

engage at this stage · Pipeline builds, watermark semantics, Auto Loader design · see the practice →

Gold rows are in the warehouse — now, what does keeping them fast cost? ↓ 02 · platform & benchmarking

02 Platform & benchmarking

We run the benchmark. You keep the harness.

The stream above just committed its rows; keeping them fast and affordable is the next problem. Compute selection, cost modeling, platform bake-offs: we design the workload, run it under controlled conditions, and hand over an instrument your team can re-run without us. The published record is the credential — including the findings that cut against the obvious answer.

  • 4,750+ TPC-DS query executions in the published record
  • 55.8B rows in the largest test environment
  • 2 × 6 platforms × compute configurations benchmarked in print
  • Yours the harness ships with the findings — re-run it after we leave

engage at this stage · Benchmarks, compute-plane selection, cost modeling · see the practice →

supporting evidence 01 · a published finding — snowflake gen2 vs gen1, tpc-ds 1tb

On TPC-DS at 1TB with small warehouses, a Gen2 warehouse finished in 0.64 minutes where Gen1 took 38.24, while consuming 9.66 credits against Gen1's 10.58.

Runtime (minutes)
Gen2 small Gen2 small: 0.64 minutes 0.64 Gen2 small: 0.64 minutes Gen1 small Gen1 small: 38.24 minutes 38.24 Gen1 small: 38.24 minutes 0 20 40
Credits consumed (credits)
Gen2 small Gen2 small: 9.66 credits 9.66 Gen2 small: 9.66 credits Gen1 small Gen1 small: 10.58 credits 10.58 Gen1 small: 10.58 credits 0 6 12

TPC-DS 1TB, cold start, small warehouses. Same size, nearly the same credits — but a 60× runtime gap driven by memory spilling. At large sizes, with no memory pressure, the result inverts and Gen1 wins the majority of queries.

source · Snowflake Gen1 vs. Gen2 vs. Snowpark-optimized warehouses · February 5, 2026

View the data as a table
MeasureWarehouseValue
RuntimeGen2 small0.64
RuntimeGen1 small38.24
Credits consumedGen2 small9.66
Credits consumedGen1 small10.58

supporting evidence 02 · the instrument — reproducibility, live

Measured and paid for — now people need to work with it. ↓ 03 · applications & bi

03 Applications & BI

Clean, cutting-edge apps — like the ones running on this page.

A measured platform still needs its human layer. The claim there is craft, and the proof is local: every interactive on this site is built the way we build client applications — hand-written, token-themed, honest to the pixel. Below, the two halves of the human layer as exhibits: reading at any scale, and writing with an audit trail.

  • 8 live canvases running on this page — demos and paintings
  • 0 frameworks, libraries, or external requests
  • 81 KB of JavaScript — all eight, total
  • <2ms per animation frame, measured
  • Faithful in both themes and under reduced motion

exhibit 01 · reading — the resampler: billions of points, live

262,144 points live demo · min–max downsampling with an adaptive lens · synthetic data

the technique from the billion-point write-up and the Mercedes engagement · move across it

exhibit 02 · writing — governed write-back: every edit an auditable transaction

from the supply-planning engagement · read the case

engage at this stage · Warehouse-backed applications and BI-layer builds · see the practice →

Humans read and write; agents need something more — context. ↓ 04 · ai agents

04 AI agents

Context is a graph, not a keyword.

Humans get dashboards; agents need context. The argument from the knowledge-graph study — and the reason we built lakecode, our own assistant with compiled, persistent context. Click any entity and compare flat retrieval with a bounded graph walk.

The platform · lakecode.ai

lakecode

Our AI assistant and agent platform, and the thing the argument above is built out of. Its context is compiled into a persistent memory substrate rather than re-fetched each session, so what it knows about your systems accumulates, survives across sessions, and is governed like any other data asset. It speaks MCP, so your team keeps whatever agent front-end it already uses.

We offer it to clients who need it — deployed against your warehouse and your codebase, rather than rebuilding the same substrate from scratch inside an engagement. Where it is not the right fit, the architecture work stands on its own and we will say so.

from the knowledge-graph study · read it

engage at this stage · Agent architecture, context engineering, and lakecode deployed on your stack · see the practice →

Every stage above moves data that has to stay protected. ↓ 05 · security & tokenization

05 Security & tokenization

Protect it at the source and it stays protected.

The layer that crosses every stage: data protection that survives the workload. Our July 2026 study argues for tokenizing before the API boundary — written as Claude Fable 5 shipped with a 30-day retention window. The follow-up benchmark, tokenization vs. masking vs. deterministic encryption across five experiments, is in progress and will publish like everything else here.

engage at this stage · Tokenization architecture and scheme selection · see the practice →

That is the stack, and every stage of it was demonstrated rather than described. ↓ — · the research

— The research

Every number above has a paper behind it.

The five stages above are demonstrations, not claims — each one runs on this page, and you can watch it work before you believe it. What stands behind them is published: 13 studies and a conference talk, every figure carrying a venue and a date, alongside the client systems those figures were measured on.

— Contact

Start with the problem.

No discovery deck, no junior on the call. A paragraph about what your platform is doing, or should be doing, is enough to begin — and we will say plainly if it isn’t work we should take.

sachin@lakesideanalytics.io

brooklyn, new york · practice established 2021