# What is DBT and why so many companies use it?

Source: https://qvantx.com/insights/what-is-dbt-and-why-so-many-companies-use-it/
Author: Ani Björkström — QvantX Sweden AB, Stockholm
Published: 2026-07-19
Updated: 2026-09-20
Video: https://www.youtube.com/watch?v=gqNMwi9YICw (6 min)
Topic: Data Engineering

License: free to quote and cite with attribution to https://qvantx.com

---
_DATA ENGINEERING_

## What is dbt, and why do so many companies use it?

#### Key takeaways

- dbt (data build tool) transforms data already in your warehouse using plain SQL, adding version control, testing, and documentation — if you know SQL, you already know 90% of dbt.

- dbt is only the T in the pipeline: it writes SELECT-based models and manages dependencies, while the warehouse — Snowflake, BigQuery, or Fabric — does the processing.

- Without dbt, your real test suite is stakeholders noticing broken numbers; with dbt, tests live in a schema YAML file and run with one dbt test command.

dbt sits between the raw data in your warehouse and the reports your stakeholders see. Ani Björkström, a tech consultant based in Stockholm, explains where dbt fits in the pipeline, what life without it looks like, and when to skip it.

### Where does dbt fit in the data pipeline?

dbt sits between the warehouse, where source data lands raw and messy, and the presentation layer — Hex, Streamlit, or Power BI reports. Upstream, sources like CRM databases feed extract-and-load tools, which save data into a warehouse such as Snowflake, BigQuery, or Fabric.

dbt itself does not process your data. It is the brain: you write models, which are just SELECT queries in SQL, and the warehouse runs them — in Snowflake, your Snowflake warehouses execute the dbt models.

### How does dbt know which tables to build first?

Through the ref function, which names the source each model reads from, dbt builds a dependency graph and knows which models depend on which. When you run dbt run, it respects that order — staging models like stage_orders and stage_customers first, then the enriched data mart, then downstream contracts like daily_revenue.

The graph also protects the team, not just the run order. Everyone can see that daily_revenue reads from orders, so a staging change immediately signals which downstream tables may need updating.

### What is life without dbt actually like?

Without dbt, transformations live in stored procedures and SQL files on a shared drive, run by scheduled jobs. The setup works — until a new joiner cannot learn the execution order, and there is no version control, testing, or documentation.

This is not because data people are sloppy; SQL alone has no framework for dependencies, tests, or docs. dbt fixes each gap: models live in Git, tests like "order_id must be unique and not null" sit in a schema YAML file and run with dbt test, and dbt docs generates a searchable site covering every table, column, and linkage from raw sources to dashboards.

### When should you use dbt — and when should you skip it?

Use dbt when transformations are SQL in a cloud warehouse, more than two people work on the project, broken data has a real cost, and you have outgrown a handful of queries. Skip it in the cases below.

| Use dbt when | Skip dbt when |
|---|---|
| Transformations are SQL in a cloud warehouse | You are still choosing an extract-and-load tool — dbt only transforms data already loaded |
| More than two people work on the project | You are the only person on the team |
| Broken data has a real cost | Your work is mainly Python and machine learning |
| You have outgrown a handful of queries | You are not comfortable with Jinja or framework conventions |

#### FAQ

**Does dbt replace my data warehouse?**

No. dbt writes and orders SQL models, but the warehouse does all the compute when the models run.

**Do I need to learn a new language for dbt?**

No. dbt models are SELECT queries in plain SQL, so knowing SQL covers 90% of dbt; the rest is conventions like the ref function and schema YAML files.

**How does testing work in dbt?**

You declare expectations in a schema YAML file and a single dbt test command runs all of them, instead of waiting for a stakeholder to report a broken dashboard.


## Full video transcript

Chapters: 0:00 Intro · 0:18 What is dbt? · 0:36 The modern data pipeline · 1:58 One function: ref() · 3:05 Life WITHOUT dbt · 4:04 Life WITH dbt · 5:10 Should YOU use it?

[0:00] [music] Hi, my name is Annie. I'm a tech consultant based in Stockholm. In this tutorial, I will tell you what is DBT and when you should and shouldn't use DBT. Let's get started. During this tutorial, we will cover what is DBT, why it matters, and when to skip it. The DBT is used to transform data already in your warehouse using plain SQL and it comes also with version control testing and documentation. DBT stands for data built tool. So if you look into the data pipeline you can see that first we have sources and those can be your CRM databases or it can be just a data providers providing data.

[0:45] Next you have your extract and load tools which can be fit or air tables for example and then after that data is saved in the warehouse and for the warehouse you can use any warehouse it can be snowflake bigquery fabric and so on and and the source data right now is saved raw and messy in your warehouse. Then after that you need to transform the data and then the end is to present this data to stakeholders via report. It can be for example hex streamlid or powerbi reports and dbt comes between warehouse where you have your source raw data and between presentation layer. So you use dbt to transform the data and dbt is only the t in the transformation step. DBT on its own doesn't process your data.

[1:35] It only writes different models and those models are just select queries that you write using SQL. And once the models are written then you run DBT models using your warehouse. For example, in Snowflake you use your Snowflake warehouses to run those DBT models. So DBT is the brain. One thing that DBT does really really well is that it is defining the order between tables because when you write DBT in DBT models you use this reference function that is saying ref and then the name of the source that is used. So DBT will write this D or dependency graph for you and it will know exactly which models are dependent on which models. In here for example you can see we have two staging models stage orders and stage customers.

[2:26] Then you have one data marked that is all just enrich. And you can see that data marked is getting data from staging models via the ref function and which means that dbt will know for example that every time if there is a change in staging then maybe orders need to be updated or the other way around then also the team knows that there is one contract that is called daily revenue and this table is reading data from orders. So dbt knows about this dependency and when you run dbt run it will respect the dependency. It will first run stages then order and reach then it will run daily revenue. Now let's try to understand what is life without dbt. You will use different store procedures on your data warehouse. You will have SQL files on a share drive or on your computer and then you will run those jobs on its own.

[3:16] They will be scheduled and they will run the procedures. So everything works fine and I worked in this setup many many times. But the issue is that if you are new and you don't know what is the order then you are lost. If you're enjoying this tutorial please give me thumbs up and comment DBT so that I continue creating these type of videos. Now back to the tutorial. Another thing is that in nonbt environment you don't have version control. You can't do testing and you can't do documentation. you know that something isn't working when a stakeholder writes you about this. So your actual test is stakeholders noticing the issues and this is not that data people are sloppy. It's just SQL alone has no framework for dependencies tests or documents and when you have the DBT you write your DBT models in your own development environment and they are saved in g and then you can push them to some type of g repository.

[4:14] So you have a version control and if you know SQL then you already know 90% of DBT. Testing is built in DBT. You write schema YAML file where for example you are writing the name of the column and then test and then under that you are saying that for example order ID shouldn't be new it should be unique. So you write this information in your schema YAML file and then when you run DBT test all of those tests are run. Same thing when it comes to documentation. So you write short description next to your models and then when you run dbt doc generation DBD builds a searchable site every table every column every linkage from row sources to dashboard and you can see all the dependencies as well in your documentations.

[5:02] So DBT isn't a new language. It is software engineering discipline wrapped around the SQL you already know. And should you actually use DBT or not? Let's decide together. the honest decision guide and use DBT. When transformations are SQL in a cloud warehouse, there are more than two people working with the project. Broken data has a real cost and you have outgrown a handful of queries. And do not use DBT if you are still finding out which extract or load tool to use because as we already saw, DBT is for transforming data when the data is already in your data warehouse. Also do not use if you are mainly working with Python, ML and you do not have more than one people in your team or if you are not convenient with Ginga or framework options.

[5:50] Want to comment DBT and I will do the next episode where I will show how to install DBT build your first model. That's it for today. Have a good day. Bye.
