— Insights & research

The measurements, in public.

Thirteen articles and a conference talk, every number dated and sourced. Published through Capital One Software, Plotly, and Databricks SME Engineering. The methodology is in the article, so you can check the work before you hire us — including the findings that cut against the obvious answer.

Analysis July 20, 2026 · Capital One Software How to use Claude Fable 5 and Mythos models with enterprise data Claude Fable 5 ships with a 30-day prompt and output retention window. The case for tokenizing at the source, before the API boundary, so protection survives the workload. Framework June 9, 2026 · Capital One Software Do I need a graph database? A framework to evaluate graph DBs When a dedicated graph database earns its place against relational primitives — and when it does not. Guide May 6, 2026 · Capital One Software Scaling agent context with knowledge graphs on Snowflake How knowledge graphs extend the working context available to AI agents beyond fixed dashboards and flat retrieval, using Cortex features and relational primitives. No dedicated graph database required. Review April 16, 2026 · Capital One Software Snowflake CoCo CLI: hands-on review for TPC-DS 10TB Five test areas, from catalog discovery through dbt model generation, run against TPC-DS at 10TB — 55.8 billion rows. Build March 19, 2026 · Capital One Software Building a GenAI cost supervisor agent in Databricks Registers 20 Unity Catalog SQL functions over Databricks system tables so an agent answers GenAI cost, governance, and attribution questions on demand — and makes the case for catalog functions over text-to-SQL to eliminate the injection surface. Benchmark February 5, 2026 · Capital One Software Snowflake Gen1 vs. Gen2 vs. Snowpark-optimized warehouses TPC-DS at 1TB across warehouse generations. On small warehouses Gen2 finished in 0.64 minutes where Gen1 took 38.24, driven by memory spilling — but at large sizes, with no memory pressure, Gen1 won the majority of queries. Benchmark January 8, 2026 · Capital One Software Jobs Classic vs. Jobs Serverless vs. DBSQL: who wins on TPC-DS? The full TPC-DS suite — 4,750+ query executions — across three Databricks compute planes via Apache JMeter. Serverless SQL warehouses won overall, faster at the tail and cheaper for the same workload, while Jobs Classic beat both Jobs Serverless variants on cost and consistency. Build March 25, 2024 · Plotly · Lead author Amplify your organization’s custom LLM strategy using Databricks with Plotly A full-stack Dash application that deploys Hugging Face and Databricks models to GPU serving endpoints. Build October 31, 2023 · DBSQL SME Engineering · Lead author Visualizing a billion points: Databricks SQL, Plotly Dash, and the Plotly Resampler At-scale interactive Dash apps over large IoT time-series via the Databricks SQL connector, with Plotly Resampler downsampling on a Polars backend. Build October 26, 2023 · Plotly · Lead author Build real-time production data apps with Databricks and Plotly Dash A Databricks Structured Streaming pipeline ingesting real-time IoT sensor data, aggregated through Auto Loader and windowing into Gold-level Delta views, served to a Dash app over Databricks SQL endpoints. Guide 2023 · Plotly · Lead author Building Plotly Dash apps on a lakehouse with Databricks SQL (advanced edition) Connecting Dash to Databricks via the SQL connector or SQLAlchemy ORM — streaming dashboards, DDL and object mapping, and advanced visuals. Guide 2023 · Plotly · Contributor Databricks SDK + Plotly Dash — the easiest way to get jobs done Dash as a front end for the Databricks SDK and Jobs API. Case study February 7, 2023 · Plotly · Contributor Molson Coors streamlines supply planning workflows with Databricks & Plotly Dash Replacing a manual spreadsheet process for tracking product ship dates with a Databricks SQL–backed Dash application — AG Grid editing with SQLAlchemy write-back, taking the workflow from 60+ steps to fewer than 10. Talk 2024 · Data + AI Summit · Speaker Petabyte Pitstops with Mercedes, Databricks SQL and Plotly Resampler A high-density visualization tool combining Rust-based downsampling, Databricks SQL, and Apache Arrow zero-copy buffers, with queries that adapt to the analyst’s zoom level.

— Contact

Want this kind of measurement on your own platform?

The method behind these articles is the one we bring to client engagements.

sachin@lakesideanalytics.io

brooklyn, new york · practice established 2021