Data Quality Checks Belong Inside the Pipeline
Data quality is often discussed as a reporting concern: profile a dataset, publish a score, and ask a steward to investigate the problems. That approach is useful for governance, but it is too late for many engineering failures. If a pipeline has already transformed and published invalid records, every downstream consumer now has to decide whether to trust the output.
Fabric supports quality work in several places, from notebooks and SQL validation to Microsoft Purview quality scans and materialized-lake-view constraints. The current DP-700 scope expects data engineers to monitor and optimize analytics solutions, and quality is part of that operational responsibility because a fast pipeline that publishes bad data is not healthy.
For Fabric Data Engineer Associate candidates, the practical rule is to move important quality checks as close as possible to the transformation that can act on them. Measure broadly, but enforce the rules that protect downstream contracts inside the pipeline.
Quality rules should describe business expectations, not generic cleanliness
Data quality includes completeness, validity, uniqueness, consistency, timeliness, and other dimensions, but not every dataset needs every metric. A customer table may require unique customer IDs and valid status values. An event stream may tolerate duplicate raw messages but require a deduplicated curated output. A financial table may need totals to reconcile to a source control amount.
Start with what would make the dataset unsafe to use. Those conditions become acceptance rules. Generic checks such as “no nulls anywhere” often create noise because some fields are legitimately optional, while important domain errors can pass because they are syntactically valid.
A good quality rule has an owner, a reason, a threshold or expected condition, and a defined response when it fails.
Profile first so thresholds come from evidence
Data profiling reveals distributions, null rates, uniqueness, ranges, and relationships that help engineers understand normal behavior. Without that baseline, quality thresholds are guesses. A rule that rejects more than one percent nulls is meaningless if the field has historically been five percent optional.
Profiling is especially important during onboarding of a new source. Inspect several representative periods, not one convenient sample. Seasonal data, end-of-month processing, and special business events can produce legitimate patterns that look anomalous in a narrow window.
Once the baseline is understood, convert stable expectations into automated checks and keep exploratory metrics for monitoring. Not every metric should block a pipeline.
Fail, quarantine, or warn based on the consequence of bad data
A quality violation can trigger different actions. A hard failure is appropriate when publishing the data would be more damaging than delaying it, such as a broken key or missing required control total. Quarantine is useful when a small number of bad rows can be isolated while valid records continue. A warning is appropriate when the condition needs investigation but does not invalidate the dataset.
The response should be designed before the rule fires. If a pipeline quarantines rows, define where they go, how they are corrected, and whether they are later merged back. If a pipeline fails, define who owns remediation and how downstream freshness is communicated.
Quality logic becomes operationally dangerous when it can stop production but nobody knows the recovery procedure.
Place checks at boundaries where responsibility changes
The most valuable quality checks often sit at handoff points: after extraction from a source, after a major transformation, before a curated table is published, and before data is exposed to critical consumers. Each boundary protects a different contract.
At ingestion, verify that the expected data arrived and can be parsed. After transformation, validate keys, relationships, and business rules introduced by the transformation. Before publication, confirm completeness, freshness, and reconciliation against expected totals.
This boundary-oriented approach fits broader data management. Ownership becomes clearer because the team can identify whether a defect originated upstream, was introduced during processing, or appeared during serving.
Row-level rules and dataset-level rules answer different questions
A row-level rule can identify an invalid date, negative quantity, malformed code, or missing required field. Dataset-level rules look at the population: row count, uniqueness ratio, distribution shift, reconciliation total, freshness, or referential coverage.
Both are necessary. Every row can satisfy its local constraints while the dataset is still wrong because half the source records were never loaded. Conversely, a dataset can have the expected row count while containing individually invalid records.
Design checks in both categories and report them differently. Row-level violations may be quarantined with error reasons; dataset-level failures often require stopping publication because the entire load is suspect.
Quality checks should preserve evidence for diagnosis
A failed rule should produce enough context to investigate without rerunning the whole workload blindly. Capture the rule name, run identifier, source, affected table, count of violations, representative samples where safe, and the expected condition. For privacy-sensitive data, store identifiers or aggregates rather than unrestricted payloads.
Historical quality results are also valuable. A gradual rise in invalid values can reveal upstream process decay before the threshold is crossed. Trend data helps distinguish a one-time source incident from a structural change that requires a new contract.
Treat quality results as operational telemetry. They are part of the evidence needed to explain why a dataset was accepted or rejected.
Do not let quality logic become an unowned second transformation layer
Complex validation can quietly duplicate business logic. If a quality rule recalculates revenue with a different formula from the transformation, disagreements become difficult to interpret. Prefer checks that validate invariants or reconcile to independent source controls rather than rebuilding the entire pipeline in parallel.
Centralize reusable domain rules when several pipelines depend on the same definition. Version them alongside transformations so an intentional business change does not create false failures. Quality code should be testable and reviewable just like production transformation code.
The goal is confidence, not a maximum number of checks. Every rule adds maintenance cost and should protect a meaningful failure mode.
Publish quality status with the data product
Consumers benefit when freshness and quality are visible rather than implicit. A curated table can expose the last successful load time, whether the latest run passed all critical checks, and which nonblocking warnings remain open. That context helps analysts distinguish a real business change from an upstream data incident.
Governance tools such as Purview can provide broader profiling and quality views, while pipeline-level checks protect immediate publishing decisions. These approaches complement each other: governance measures the estate, and engineering controls prevent known defects from moving through critical paths.
Quality becomes trustworthy when it is continuous, actionable, and connected to ownership. Put the decisive checks inside the pipeline so bad data does not have to reach a dashboard before someone notices.
Use quality results to improve the source contract
Quality controls should not become a permanent downstream cleaning service for avoidable source defects. When a rule repeatedly catches the same problem, feed that evidence back to the producing system. If customer identifiers are malformed every week, the most valuable fix may be validation in the source application. If a partner file consistently arrives late, the issue is a delivery contract rather than a transformation problem.
Track recurring violations by source, rule, and business owner. That history helps distinguish one-off operational noise from structural quality debt. It also gives governance teams a way to prioritize improvements according to actual downstream impact instead of subjective complaints. A small set of recurring defects often accounts for a large share of remediation effort.
Be careful with automatic repair. Standardizing known formats, trimming whitespace, or mapping approved reference values can be appropriate, but silent correction can hide upstream problems and sometimes changes meaning. Preserve the original value or a traceable audit record when the transformation makes a material correction. Consumers should be able to explain how a published value was derived.
Validation strategy should also match the cost of recovery. Cheap deterministic checks can run on every batch, while expensive statistical tests may be better on representative samples or scheduled full scans. The important requirement is that a failed check has a clear path to correction and replay. When bad records are fixed upstream or a rule is corrected, engineers should know which partitions, files, or load windows must be reprocessed and how to prove that the replacement output is complete. Keep the failed run and its evidence long enough to compare it with the successful replay. That comparison turns quality enforcement into an auditable recovery process instead of a binary red-or-green status that loses the history of what actually happened.
The mature end state is not a pipeline with thousands of checks. It is a data supply chain in which producers understand their contracts, transformations validate what they change, and consumers receive clear evidence about freshness and trustworthiness.
Quality thresholds should also be reviewed when the business changes. A sudden rise in nulls can indicate a defect, but it can also follow an intentional process change that makes a field optional. A new product line can introduce valid categories outside yesterday’s reference list. If engineers respond by simply loosening every failing rule, controls lose value; if they refuse to update rules, the pipeline becomes a source of false incidents. Treat rule changes as governed changes: document the reason, confirm the business owner, test historical impact, and update both the transformation and the quality expectation when necessary. This keeps quality checks aligned with meaning rather than freezing the data model at the moment the rule was first written. Strong quality systems evolve, but they evolve deliberately.