
Defining architectural points for data quality control
In complex integration landscapes, several key points exist where data quality checks can be embedded. These points determine how early problems are detected and their impact on subsequent processes. We distinguish between source-level, in-transit, transformation-level, and target-level checks.
- Source-level validation: Performed directly in the source system or during data extraction. This allows problems to be identified before data enters the integration flow.
- In-transit validation: Embedded in integration buses, message brokers, or Data Quality Gates[1], checking data “in motion” in real-time[1].
- Transformation-level validation: Applied during data transformation, for example, in ETL processes, where transformation occurs before loading into the target system to ensure high data quality at early stages[5].
- Target-level validation: Performed after data is loaded into the receiving system, typical for ELT architectures where raw data is loaded directly, and transformation happens there and shifts the responsibility for quality to later stages[5].
Advantages and disadvantages of placing data quality checks at different stages
Choosing the location for data quality checks is a compromise between early detection, performance, and cost. Early detection of data problems, known as “shift-left testing,” significantly reduces the cost of error correction[2], as fixing them at later stages is much more expensive. For example, poor data quality can lead to significant financial losses, estimated at an average of $15 million annually for individual organizations[2].
- Early checks (Source-level, In-transit):
- Advantages: Rapid detection and prevention of poor-quality data propagation. Reduced cost of error correction.
- Disadvantages: May require intervention in source systems. Additional load on source systems or integration buses. Complexity of implementation for a large number of heterogeneous sources.
- Late checks (Transformation-level, Target-level):
- Advantages: Allow for more comprehensive checks after data aggregation or transformation. Less impact on source system performance.
- Disadvantages: Higher error correction costs as they are detected later. Risk of poor-quality data spreading through the system before detection. Can slow down the process for large data volumes in ETL architectures (ETL)[5].
Strategic considerations for integrating data quality control
When choosing the optimal location for data quality checks, architects must consider several key factors:
- Compliance and regulatory requirements: Some industries have strict data quality requirements, which may dictate the need for early and comprehensive checks.
- Data volume and velocity: For streaming data, real-time validation is necessary to detect problems before they spread through data pipelines and can lead to significant financial losses[2]. Batch processing allows for comprehensive checks of historical data (streaming data)[2].
- Latency requirements: Systems requiring low latency may necessitate minimal checks in the integration flow, shifting more complex validations to asynchronous processes.
- Autonomy of source and target systems: The degree of control over source systems influences the feasibility of implementing checks at early stages.
- Available tools and technologies: Utilizing centralized data quality services or Data Quality Gates[1] can simplify data quality management.
- Impact on overall Data Governance strategy: Decisions regarding check placement must align with the organization's overall data governance principles.
Architectural patterns for embedding data quality control
Various architectural patterns are applied for effective data quality control in complex integration flows:
- Centralized data quality services: Creating a separate service responsible for executing all data quality rules. This centralizes logic, simplifies management, and promotes rule reuse.
- Embedded checks in API Gateways or message brokers: Using an API Gateway or message brokers (e.g., Kafka) to perform basic data validity and format checks before further processing.
- Data Quality Firewalls[1] or Data Quality Gates[1]: These patterns involve creating checkpoints through which data passes, where sets of quality rules are applied. They allow for stopping or redirecting poor-quality data, preventing its spread.
- Patterns for streaming and batch processing: For streaming data, the focus is on lightweight, real-time checks, while for batch processes, more in-depth and resource-intensive analyses are possible. Modern architectures, such as Data Fabric and Data Mesh, aim to automate data discovery, management, quality control, and problem monitoring closer to the source (Data Fabric and Data Mesh).
Failure mechanism: Propagation of poor-quality data
Insufficient or improperly placed data quality control can lead to the propagation of poor-quality data throughout the system. This can result in incorrect business decisions, loss of trust in data, operational failures, and significant financial losses. Poor data quality can cost the U.S. economy $3.1 trillion[2].
DMIG, as a company specializing in system integration and data management, understands the critical importance of strategically placing data quality controls. Our architects use these principles to create robust integration solutions that minimize the risks associated with poor-quality data and optimize error correction costs for our clients.
Decision matrix for choosing data quality control placement
This matrix helps architects evaluate different architectural patterns for data quality control based on key criteria. For each criterion, assess the impact or characteristic for each placement option (e.g., low, medium, high).
| Criterion | Source-level | In-transit | Transformation-level | Target-level |
|---|---|---|---|---|
| Cost of error correction | Low | Low/Medium | Medium | High |
| Impact on performance | Medium/High (on source) | Medium | Medium/High (on ETL) | Low (on ingestion) |
| Implementation complexity | Medium/High | Medium | Medium | Low/Medium |
| Time to detect errors | Immediate | Near real-time | After extraction/before loading | After loading |
| Data volume | Any | Medium/High (with optimization) | High (batch processing) | High (raw data) |
| Latency requirements | Low | High (for real-time) | Medium | Low |
| Level of trust in source | Low | Medium | Medium/High | High |
| Compliance requirements | High | High | Medium | Medium |
How to apply the matrix:
- Prioritize: For your specific project or integration flow, identify which criteria are most important (e.g., low error correction cost, high performance, minimal latency).
- Evaluate each option: Analyze each check placement option (source-level, in-transit, transformation-level, target-level) against your prioritized criteria, using the ratings from the table.
- Weigh trade-offs: No single option will be ideal for all criteria. Use the matrix to identify trade-offs and justify your choice. For example, if early detection and low correction cost are critical, prioritize source-level or in-transit checks, despite potential impact on source performance.
- Combine approaches: Often, the optimal solution is a hybrid approach, combining basic checks at early stages with more comprehensive validations at later stages.
Strategic placement of data quality control is not just a technical task but a key architectural decision that impacts data reliability, business process efficiency, and the overall TCO / ROI of integration solutions. Careful analysis and a well-reasoned choice of check placement are crucial for success in complex integration landscapes.
Перелік джерел

Author
