— Services

Five practices. Each ends in something running.

Architecture, benchmarking, applications, AI agents, and data security. Engagements are scoped, time-boxed, and handed over. Whether the deliverable is a benchmark report, a migration plan, an application, an agent, or a tokenization design, your team keeps something they can run and maintain without us.

01 · Architecture

Databricks & Snowflake platform architecture

Platform design and remediation for teams already committed to a lakehouse or cloud warehouse — and for teams deciding between them. We work on the structure that determines cost and speed for years: compute topology, storage layout, governance model, and the migration path between them.

  • Lakehouse and warehouse design reviews, with a written findings document
  • Unity Catalog rollout: governance model, permissions, catalog structure
  • Compute topology — warehouse sizing, cluster policy, workload isolation
  • Migration planning between platforms or between compute planes
  • Query and pipeline tuning against a measured baseline

02 · Benchmarking

Compute benchmarking & cost optimization

Independent measurement of what your platform actually costs and how fast it actually is. We build reproducible harnesses over standard suites and over your own workloads, so the numbers survive scrutiny and can be re-run after the engagement ends. This is the practice the rest of the work is built on.

  • TPC-DS and custom workload benchmarking with JMeter-driven harnesses
  • Compute plane comparisons — serverless vs. classic vs. SQL warehouse
  • Spend analysis from system tables and account usage views
  • Cost-per-query and price/performance modeling at your concurrency
  • A harness your team owns, so results can be reproduced later

03 · Applications

Custom data & analytics applications

Data products for the cases a BI tool cannot reach: very large time-series, real-time streams, custom interaction models, or performance requirements that need work below the framework. We build full-stack — warehouse connection through interface — and drop into Rust and Arrow when the performance ceiling is the actual constraint.

  • Interactive analytics apps over Databricks SQL and Snowflake
  • Billion-point time-series visualization with server-side downsampling
  • Real-time and streaming dashboards
  • High-performance data engines in Rust, Arrow, and Polars
  • Model-serving front ends and internal analytical tooling

04 · AI agents

AI agents & GenAI on the data platform

Agents that are actually grounded in your data and actually governed. The hard parts are context and safety: giving an agent more working context than flat retrieval provides, and exposing your warehouse through a tool surface that cannot be talked into arbitrary SQL. We build lakecode, our own assistant and agent platform, whose context is compiled into a persistent memory substrate rather than re-fetched each session — and where it fits your problem, we deploy it for you rather than rebuilding the same substrate from scratch. Where it does not, the architecture work stands on its own. And we make GenAI spend legible before it becomes a line item nobody can explain.

  • lakecode, deployed on your stack: an assistant with compiled, persistent context, over MCP so your team keeps its own front-end
  • Agent memory and context engineering — the substrate work behind lakecode
  • Knowledge-graph context so agents reason past fixed dashboards
  • Governed tool surfaces — catalog functions instead of open text-to-SQL
  • MCP server design and integration against internal data systems
  • GenAI cost supervision built on platform system tables
  • Evaluation harnesses so agent quality is measured, not assumed

05 · Data security

Tokenization & data protection for AI

Protection applied at the source survives every workload downstream — that is the argument of our July 2026 study on running frontier models against enterprise data, written as Claude Fable 5 shipped with a 30-day retention window. We work on tokenize-before-the-API-boundary architectures and on choosing the right scheme per workload; the follow-up benchmark, tokenization against masking against deterministic encryption across five experiments, is in progress and will publish like everything else here.

  • Tokenize-at-source architectures ahead of LLM and agent API boundaries
  • Scheme selection — tokenization vs. masking vs. deterministic encryption, by workload
  • Protection–utility analysis: what survives analytics, ML, and agent use
  • PII surface review across pipeline, warehouse, and AI layers
  • Retention-term review for hosted-model integrations

02 Engagement model

Three ways to start.

Two to four weeks

Assessment

A focused review — architecture, spend, or a specific workload — producing a written findings document with measured evidence and a prioritized set of recommendations.

Fixed scope

Project

A scoped build: a benchmark harness, a migration, an application, or an agent. Agreed acceptance criteria, documentation and handover included.

Month to month

Advisory

Recurring time for teams who want a second set of eyes on architecture decisions, vendor claims, and performance work as they come up.

— Contact

Describe the problem — that’s enough to start.

A paragraph about what your platform is doing, or should be doing, is plenty for a first conversation. We will tell you honestly if it isn’t work we should take.

sachin@lakesideanalytics.io

brooklyn, new york · practice established 2021