— Work

Systems built and measured at production scale.

Three engagements with public write-ups, and the tools built alongside them. Clients are named only where the work is already public through a talk title, an article title, or a byline.

Trillions of rows, explorable live
Mercedes Petabyte time-series visualization 2023–2024

Trillions of rows, explorable in real time.

The analytics team needed to explore time-series data at a scale where the usual answer — pre-aggregate everything and accept the loss of resolution — destroyed the signal they were looking for.

We built a high-density visualization tool combining Rust-based downsampling with Databricks SQL and Apache Arrow zero-copy memory buffers. Queries adapt to the analyst’s zoom level: aggregated views by default, with SQL pushdown retrieving original high-resolution data on demand. Larger-than-memory datasets stay interactive instead of being flattened into something smaller and less useful.

“Petabyte Pitstops with Mercedes, Databricks SQL and Plotly Resampler” · Data + AI Summit 2024

Trillions of rows handled without pre-aggregation
Zero-copy Arrow buffers for larger-than-memory datasets
DAIS 2024 presented at Data + AI Summit
Remembers context compiled, not re-fetched
lakecode AI coding assistant · our own product 2025–present

An assistant that remembers the codebase.

Consulting on agents taught us where they fail: context. Retrieval fetches fragments per session and forgets; the working knowledge an agent needs is a structure, not a search result.

So we built lakecode — an AI coding assistant whose context is compiled into a persistent memory substrate instead of re-fetched every session. The same context-engineering argument as the knowledge-graph study, applied to code: what the agent knows accumulates, survives across sessions, and is governed like any other data asset. Exposed over MCP, so teams keep their own agent front-ends.

the product, public · lakecode.ai

Compiled context substrate, not per-session retrieval
Persistent memory that survives across sessions
MCP native tool surface — bring your own agent
4,750+ TPC-DS query executions
Capital One Software Published benchmark program Ongoing since Jan 2026

Full workloads, at scale, with the method published.

An ongoing research program measuring compute economics and agent architecture across Databricks and Snowflake — the questions platform teams actually decide on, answered with standard suites run under controlled conditions.

Every study ships with the notebooks, worksheets, or prompt lists needed to check the work, including the findings that cut against the obvious answer. When Gen2 warehouses lost to Gen1 at large sizes, that went in the piece.

seven studies, Jan–Jul 2026 · read the research

4,750+ TPC-DS query executions across three compute planes
55.8B rows in the largest test environment, at 10TB
7 studies published, Jan–Jul 2026
60+ → <10 workflow steps
Molson Coors Supply planning on Databricks SQL 2023

Planners write to the warehouse, not to a spreadsheet.

The supply planning team tracked product ship dates through a manual spreadsheet process — extracts passed between people, with no auditable record of what changed or why.

We replaced it with a Databricks SQL–backed application: an editable AG Grid with SQLAlchemy ORM write-back, so planners keep the workflow they trust while the numbers live somewhere auditable. The published write-up puts the workflow at 60+ steps before and fewer than 10 after.

written up by Plotly · feb 2023

60+ → <10 steps in the supply planning workflow, as published
Write-back native ORM writes to the warehouse from the app

02 Also built

Tools and applications without public write-ups.

Listed plainly, without metrics — there is no published source to cite for these.

DbxOauth

An extension letting Dash applications authenticate against Databricks APIs over OAuth — both M2M service principals and U2M user-level requests — with automatic encryption, storage, and refresh of tokens.

Databricks Delta Optimizer

An application for optimizing Delta tables from user-defined strategies, integrating the Jobs API, Databricks SQL connector, SDK, and OAuth to manage authentication and scheduled runs.

Custom LLM management app

A Dash interface for non-technical users to select, register, and deploy custom and Hugging Face models to Databricks serving endpoints, with an interactive chat surface for querying them.

Real-time IoT streaming pipeline

A Databricks Structured Streaming pipeline ingesting sensor data from an Azure IoT Hub into ADLS, using Auto Loader, watermarking, and windowing to build Gold-level Delta views.

— Contact

Start with the problem.

No discovery deck, no junior on the call. A paragraph about what your platform is doing, or should be doing, is enough to begin — and we will say plainly if it isn’t work we should take.

sachin@lakesideanalytics.io

brooklyn, new york · practice established 2021