Skip to content

What is Databricks used for? A plain-English guide

Databricks is where companies bring their data together, clean it, and use it for reporting, machine learning and AI. Here is what it does, what it is not for, and how to tell whether you need it.

Bright modern manufacturing plant with a production line and a wall screen showing data charts
Contents
  1. What is Databricks used for? Five real use cases
  2. The lakehouse and medallion layers, explained simply
  3. How Databricks sits next to your ERP
  4. Manufacturing and defence: where it earns its place
  5. What Databricks is not good for
  6. When Databricks fits, and when a simpler warehouse is enough
  7. A decision checklist
  8. Where to start
  9. Frequently asked questions

Databricks is used to collect a company's data in one place, clean it up, and turn it into reports, forecasts and AI applications. In business terms, it is the platform where data from your ERP, plants, CRM and files gets engineered into something people and models can trust.

The longer answer depends on your problem, because Databricks is broad. It covers data pipelines, a SQL data warehouse, BI dashboards, machine learning, AI agents and now a Postgres database. Very few companies need all of it on day one.

I have worked with Databricks on data platform programmes in manufacturing and defence. This guide is for the CIO, CFO or head of data deciding whether it belongs in their architecture.

Databricks' own platform overview lists a long set of jobs it can do. In practice, I see five.

  1. Data engineering and ETL. Pulling data from source systems, cleaning it and loading it into tables. This is the workhorse use case and usually where most of the spend goes. The product family is Lakeflow. Lakeflow Connect handles ingestion and Lakeflow pipelines (formerly Delta Live Tables) handle transformation. Lakeflow Designer is visual data preparation, and Lakeflow Jobs is orchestration.
  2. Data warehousing with Databricks SQL. Databricks SQL is its cloud data warehouse, running on SQL warehouses. Analysts write SQL against the same tables the engineers built. No separate copy in another warehouse.
  3. BI and self-service analytics. AI/BI dashboards sit inside the platform, and the tables also feed Power BI, Tableau or whatever your business already uses. Genie lets business users ask questions in plain language against governed data.
  4. Machine learning. Demand forecasting, predictive maintenance, quality prediction, pricing. MLflow tracks experiments and models, and model serving puts them into production.
  5. Generative AI and agents. Agent Bricks is the current home for building and governing AI agents on your data. Genie One, launched as generally available in June 2026, is the business-user front end. Older material calls much of this AI tooling Mosaic AI; current product pages lead with Agent Bricks and Genie.

Two newer pieces matter for some buyers. Lakebase is a managed Postgres database inside the platform, for applications that need fast transactional reads and writes next to analytical data. Databricks announced Lakebase as generally available in early 2026. Databricks Apps, generally available since May 2025, lets teams build and host internal data and AI applications on the platform.

One naming note, because it confuses search results. Databricks spent two years marketing the "Data Intelligence Platform". Its product pages now say "Databricks Data + AI Platform". Same platform. And Databricks One was renamed Genie One.

Before Databricks, most companies ran two systems. A data lake held cheap raw files, including sensor logs, documents and images, for data scientists. A data warehouse held clean, structured tables for finance and BI. Data was copied between them, and the two copies drifted.

The lakehouse is Databricks' answer: keep one copy of the data in open formats (Delta Lake tables on your cloud storage) and run both BI and AI on it. Unity Catalog sits over the top and governs who can see what, tracks lineage, and logs access. That governance layer is what turns a pile of files into something an auditor will accept.

Most teams organise the lakehouse with the medallion architecture. Three layers, each cleaner than the last.

How data moves through the medallion layers

  1. Source systems

    ERP, MES, historians, CRM, spreadsheets and documents.

  2. Bronze

    Raw data landed as it arrived. Kept for audit and reprocessing.

  3. Silver

    Cleaned, deduplicated and validated. One trusted version of each record.

  4. Gold

    Modelled for the business: margin by product, downtime by line, spend by supplier.

  5. Consumers

    Dashboards, finance reporting, ML models, AI agents and apps.

The business leader's version: bronze is the evidence, silver is the cleaned record, gold is the answer to a business question. Your CFO should only ever see gold. Your data scientists will spend most of their time in silver.

In the manufacturing programmes I have worked on, silver is where the real effort sits. Material numbers that differ between plants. Units of measure that do not reconcile. Sensor timestamps in local time against ERP postings in UTC. None of that is a Databricks problem, but Databricks is where you end up solving it.

Databricks does not replace SAP S/4HANA, Oracle Fusion Cloud ERP or Microsoft Dynamics 365. The ERP stays the system of record for orders, stock, invoices and the ledger. Databricks takes copies of that data, combines it with data the ERP never sees, and serves analytics and AI.

The ERP data usually arrives through one of three routes:

  1. Native or partner connectors. Lakeflow Connect, plus third-party replication tools, pull tables or change data from the ERP database or its APIs.
  2. Extracts from the ERP's own data layer. Exports from SAP, Oracle or Dynamics analytics and data services, landed into bronze.
  3. The ERP vendor's data product. For SAP shops, SAP Business Data Cloud is one option. It includes SAP Databricks, a managed Databricks service inside SAP's offering, which Databricks announced as generally available. There is also a zero-copy sharing route from Business Data Cloud to an existing Databricks estate.

Which route fits depends on licensing, data volumes and how much non-ERP data matters. The hard part is never the extract. It is the business meaning: which plant, which ledger, which version of a customer. If you have not sorted that out, my piece on why data migrations fail covers the same root causes.

Manufacturing. The strongest case is joining ERP data with plant data. Machine and sensor readings from historians and IoT platforms, manufacturing execution system (MES) records, quality inspection results and maintenance logs. Joined to ERP orders, materials and costs, that data answers questions finance and operations both care about: scrap cost by line, maintenance before a breakdown, supplier quality against price paid. Sensor volumes are large and messy. A lakehouse handles that better than a traditional warehouse. In my manufacturing work, the first use case that paid for the platform was costing: bringing raw material and labour costs together to arrive at a standard cost. Finance already owns that number, so the value is easy to see. If you are still choosing the ERP underneath, see my comparison of ERP for manufacturing.

Defence. The questions are different. Where does the data physically sit? Who operates the platform? Can classified and unclassified data be kept apart, with access proved in an audit? Programmes run for decades, so the platform choice has to outlive several technology cycles. On my defence work, keeping classified and unclassified data apart shaped the design more than anything else. Settle that separation before you design a single pipeline.

Architecture matters here. In Databricks, the control plane runs in Databricks' account. Classic compute runs in your own cloud account. Serverless compute runs in Databricks' account, in the same region. For a defence data lead, that difference is the first conversation, before any feature demo. For US government work there is a separate Databricks on AWS GovCloud deployment. Outside the US, check region and sovereign-cloud availability against your residency and export-control rules (ITAR-type obligations included) before you shortlist. I have found that Unity Catalog's fine-grained access control and audit logs carry a lot of weight in these reviews.

It is easy to buy Databricks for the wrong job.

  1. It is not an ERP or a system of record. Do not run transactions or the ledger through it. Lakebase handles application data, not your general ledger.
  2. It is not a quick dashboard tool. If the goal is twenty finance dashboards on clean ERP data, Databricks is a heavy way to get there.
  3. It does not fix poor data ownership. If no one owns the customer master, Databricks will just process the mess faster.
  4. It does not run itself. You need data engineers who can write SQL and Python, plus someone watching costs. Spend is metered on usage, so badly written jobs show up on the invoice.
  5. It is not free for business use. The Free Edition is for learning and experiments. Databricks' Free Edition limitations state it may not be used for commercial purposes.

Databricks will not fix a data problem you have not defined. It will just process it faster.

The question I hear most: do we need Databricks, or would a simpler warehouse do? This is how I frame it.

Your situationDatabricks fitsA simpler option is likely enough
Data sourcesERP plus sensor, MES, documents, external feedsMostly ERP and CRM tables
Data volume and changeHigh volume, streaming or near real-timeDaily batch, moderate volume
Main usersData engineers, data scientists and analystsFinance and business analysts
AI and ML plansFunded use cases with ownersIdeas without a sponsor yet
Team skillsSQL and Python engineers in house or contractedMostly BI and Excel skills
Governance needsFine-grained access, lineage, audit across many sourcesStandard role-based access

If most of your answers sit in the right-hand column, start with your ERP vendor's analytics or a managed warehouse. For a closer comparison, read my sibling pieces on Snowflake vs Databricks and Databricks vs Microsoft Fabric. For cost, see Databricks pricing explained. Databricks bills compute in Databricks Units (DBUs), on pay-as-you-go or committed-use terms, and the cloud storage and networking sit on top.

A decision checklist

  1. Can you name three business decisions that will change if the data platform works? If not, stop and find them.
  2. Which data sources do you need beyond the ERP, and how big are they?
  3. Who owns master data for customer, material, supplier and chart of accounts?
  4. What are your data residency, sovereignty and export-control constraints?
  5. Do you have, or will you fund, a data engineering team for the long term?
  6. Which BI tool stays, and will it read from Databricks SQL?
  7. Who is accountable for monthly platform spend, and what alert triggers a review?
  8. What will you switch off when this goes live? If the answer is nothing, you are adding cost, not consolidating.

Pick one use case with a named business owner and real money attached. In manufacturing that is often scrap, downtime or inventory. Build it through bronze, silver and gold with Unity Catalog governance switched on from the first table. Measure the result and the run cost. Then decide whether to scale.

I would not start with an enterprise-wide migration or an AI agent. Governance belongs in that first build, not in a later phase; my AI governance framework guide explains why.

If you are weighing Databricks against other platforms, or working out how it sits next to SAP, Oracle or Dynamics 365, I help organisations with these decisions. You can get in touch here.

What is Databricks in simple terms?

Databricks is a cloud platform where companies store, clean and analyse large amounts of data, and build machine learning and AI on top of it. It runs on AWS, Microsoft Azure and Google Cloud. Think of it as the engineering workshop between your source systems and your reports, models and AI tools.

Is Databricks a database or a data warehouse?

Both, partly. Databricks SQL is a cloud data warehouse built on the lakehouse, and Lakebase is a managed Postgres database for applications. Underneath, data sits as open-format tables on cloud storage. It is not a transactional system of record like an ERP.

What is the difference between Databricks and Snowflake?

Both run SQL analytics on cloud data. Databricks started from data engineering and machine learning on open formats. Snowflake started as a managed SQL data warehouse. The two now overlap heavily. The right choice depends on your data types, skills and AI plans. My Snowflake vs Databricks comparison covers it in detail.

Does Databricks replace SAP or other ERP systems?

No. The ERP remains the system of record for transactions and the ledger. Databricks takes copies of ERP data, combines them with other sources such as plant and sensor data, and serves analytics and AI. SAP customers can also use SAP Databricks inside SAP Business Data Cloud.

What is medallion architecture in Databricks?

A way of organising data in three layers. Bronze holds raw data as it arrived. Silver holds cleaned and validated records. Gold holds business-ready tables for reporting, ML and AI. Each layer is more refined than the last, and business users should only consume gold.

Is Databricks free to use?

Databricks Free Edition is free for learning and experimenting, with quotas and feature limits, and may not be used for commercial purposes. Business use is billed on consumption in DBUs, either pay-as-you-go or through a committed-use contract, with cloud storage and networking on top.

Noel D'Costa

Written by

Noel D'Costa

25 years across SAP and Oracle ERP programmes in aviation, government, finance, retail, and manufacturing. Finance background. I help leadership teams scope transformations honestly, recover programmes in trouble, and build systems that survive their first year in production.

Next step

Running an ERP programme right now?

If this article touched on a programme you are live in right now, a 30-minute conversation usually gets further than another week of internal analysis.