Great Expectations (GX) is an open-source Python library for data quality testing. You describe what good data looks like, such as "order_id is never null" or "amount is between 0 and 100,000", and GX validates your data against those expectations and produces a readable report. This tutorial uses GX Core 1.x.
Note: our Udemy course does not teach Great Expectations. It teaches the Python and SQL validation skills this library builds on.
What Is Great Expectations?
The main concepts are:
| Concept | Meaning |
|---|---|
| Expectation | One assertion about data, e.g. ExpectColumnValuesToNotBeNull |
| Expectation suite | A named group of expectations for one dataset |
| Data source / asset / batch | Where data comes from (pandas, Spark, a SQL database) and which slice to validate |
| Validation definition | Links a suite to the data it should run against |
| Checkpoint | Runs one or more validations and triggers actions, such as updating Data Docs or sending alerts |
| Data Docs | HTML reports of expectations and validation results |
GX 1.x changed the API significantly from the older 0.x releases, so many older blog posts and videos no longer match. Check the version before following any tutorial.
Install
python -m venv .venv source .venv/bin/activate # Windows: .venv\Scripts\activate pip install great_expectations pandas
Your First Expectation
import great_expectations as gx import pandas as pd df = pd.read_csv("orders.csv") context = gx.get_context() # ephemeral context, fine for learning source = context.data_sources.add_pandas("orders_source") asset = source.add_dataframe_asset(name="orders") batch_def = asset.add_batch_definition_whole_dataframe("orders_batch") batch = batch_def.get_batch(batch_parameters={"dataframe": df}) expectation = gx.expectations.ExpectColumnValuesToNotBeNull(column="order_id") result = batch.validate(expectation) print(result.success) # True or False
The result also includes details such as how many values failed and a sample of unexpected values, which is exactly what you need for a defect report.
Expectation Suites
Group the checks for a table into a suite:
suite = context.suites.add(gx.ExpectationSuite(name="orders_suite")) suite.add_expectation(gx.expectations.ExpectColumnValuesToBeUnique(column="order_id")) suite.add_expectation(gx.expectations.ExpectColumnValuesToNotBeNull(column="customer_id")) suite.add_expectation(gx.expectations.ExpectColumnValuesToBeBetween( column="amount", min_value=0, max_value=100000)) suite.add_expectation(gx.expectations.ExpectColumnValuesToBeInSet( column="status", value_set=["PLACED", "SHIPPED", "DELIVERED", "CANCELLED"])) suite.add_expectation(gx.expectations.ExpectColumnValuesToMatchRegex( column="email", regex=r"^[^@]+@[^@]+\.[^@]+$"))
Validation Definitions
validation = context.validation_definitions.add(
gx.ValidationDefinition(name="orders_validation", data=batch_def, suite=suite)
)
results = validation.run(batch_parameters={"dataframe": df})
print(results.success)
for r in results.results:
if not r.success:
print(r.expectation_config.type, r.result.get("unexpected_count"))
Checkpoints
A checkpoint runs validations and actions together, which is how GX is usually wired into a pipeline step or a scheduler:
checkpoint = context.checkpoints.add(
gx.Checkpoint(
name="orders_checkpoint",
validation_definitions=[validation],
actions=[gx.checkpoint.UpdateDataDocsAction(name="update_docs")],
)
)
run = checkpoint.run(batch_parameters={"dataframe": df})
if not run.success:
raise SystemExit("Data quality checks failed") # stop the pipeline
For a persistent project with saved suites and Data Docs, create a file context with gx.get_context(mode="file").
Classic ETL Checks in Great Expectations
| ETL check | Expectation |
|---|---|
| Null check | ExpectColumnValuesToNotBeNull |
| Duplicate key check | ExpectColumnValuesToBeUnique / ExpectCompoundColumnsToBeUnique |
| Valid code list | ExpectColumnValuesToBeInSet |
| Range check | ExpectColumnValuesToBeBetween |
| Format check | ExpectColumnValuesToMatchRegex |
| Expected row count | ExpectTableRowCountToEqual / ExpectTableRowCountToBeBetween |
| Schema / column list | ExpectTableColumnsToMatchOrderedList |
One gap to know: expectations validate a single dataset. Source-to-target comparisons (missing records, transformation rules across two tables) are usually done with SQL or pandas instead. See the Python ETL testing framework for that side.
When to Use Great Expectations
- Good fit: recurring data quality checks on pipeline outputs, monitoring production data, teams that want readable reports for non-technical stakeholders
- Less suited: one-off source-to-target reconciliation, complex multi-table business rules
- Alternatives: dbt tests if you use dbt (see the dbt testing tutorial), plain SQL, or commercial tools like QuerySurge
ETL testing tool tutorials: ETL testing SQL queries · QuerySurge ETL testing · Informatica ETL testing · dbt testing tutorial · ETL testing tools · Python ETL testing framework
Frequently Asked Questions
What is Great Expectations used for?
Great Expectations is an open-source Python library for data quality testing. You define expectations about your data, validate datasets against them and get readable reports, often as a step in a data pipeline.
Is Great Expectations free?
Yes. GX Core is open source and free. There is also a paid hosted product, GX Cloud, which adds a web interface and managed scheduling.
Can Great Expectations do source-to-target testing?
Not directly. Expectations validate one dataset at a time. Source-to-target comparisons are usually written in SQL or pandas, while GX handles rules on each dataset such as nulls, uniqueness, ranges and formats.
What is the difference between Great Expectations and dbt tests?
dbt tests are SQL queries that run inside a dbt project against warehouse models. Great Expectations is a Python library that can validate pandas, Spark and SQL data, independent of dbt.
Asim Noaman Lodhi
Certified Google Partner · QA Consultant · 12+ Years IT
QA consultant specializing in ETL testing and data quality. Trained 913+ students to transition into data testing roles through hands-on, real-world instruction.