
The Cost Compass: AI in Cloud FinOps and Usage Optimization
AI can help finance and engineering teams detect waste, forecast cloud spend, and connect infrastructure usage to product value.
Read MoreZharfAI Team

Data observability tells a team what changed, where, and when. Data quality asks whether information is fit for a defined use. Neither automatically establishes that a value means the right thing. A perfectly fresh, complete table can still encode the wrong business definition.
AI can detect unusual distributions, correlate incidents with lineage, summarize failures, and suggest tests. It should not silently repair production data or declare semantic truth without accountable domain review.
Teams often collapse different questions into one “quality score.” Keep them separate:
An order-status column may pass all technical checks while mapping “cancelled after shipment” differently in two systems. Revenue can reconcile to finance and still be inappropriate for a tax filing because recognition rules differ. State the use and accountable owner before calling data “good.”
ISO 8000-8:2015, confirmed in 2022 and current as of July 2026, describes fundamental information and data-quality concepts and prerequisites for measurement within quality management. It does not prescribe one magic dashboard score.
ISO/TS 8000-81:2021, confirmed in 2025, specifies a profiling procedure for structured data, including structure, column, and relationship analysis. Its scope does not include deriving data rules or measuring the extent of nonconformity. Profile first, then define rules with domain owners.
Build dimensions that match the use:
Report dimensions separately. A composite score can hide a critical failure behind healthy columns.
A useful contract says more than column type. It records:
For a net_revenue metric, define refunds, tax, discounts, currency conversion, recognition date, test accounts, late adjustments, and restatement policy. If two teams need different definitions, publish two named metrics rather than forcing an ambiguous shared field.
Instrumentation should cover source extraction, transport, transformation, orchestration, storage, semantic models, reports, features, retrieval indexes, and exports. For each run capture code or query version, parameters, input and output identifiers, row counts, timing, status, and ownership.
OpenLineage is an open framework and extensible specification for lineage metadata built around datasets, jobs, runs, and events. It can improve interoperability in lineage collection. It does not prove that every transformation is captured or that the SQL implements the intended definition.
Test lineage completeness with known paths. Manual spreadsheets, reverse ETL, dashboard calculations, notebooks, exports, and vendor transformations often sit outside the graph. Label inferred lineage separately from runtime-observed lineage.
Static tests are best for known invariants: primary-key uniqueness, nonnegative quantity, accepted currency, mandatory consent state. Statistical detection helps with unknown change: row-count shifts, distribution movement, seasonality breaks, category appearance, and relationship changes.
For each alert define:
Tune on incidents, not alert counts. Measure precision, missed known incidents, time to detect, time to acknowledge, time to contain, recurrence, and engineer effort. A system that fires daily without ownership trains people to ignore it.
An AI incident assistant can retrieve recent deployments, failed tests, schema changes, lineage, logs, and prior runbooks. It can draft a timeline and rank hypotheses. Require citations to the underlying evidence and separate:
Never allow a model to rewrite data, change a contract, suppress an alert, or backfill history without explicit authority and reversible execution. Prompt injection can enter through logs, field values, documentation, or tickets; retrieved text is data, not trusted instruction.
Observability can show that customer_status changed from five values to six. It cannot decide whether “paused” belongs inside “active” for a regulatory report, customer-success dashboard, or invoice process.
Use a semantic review packet:
Maintain a metric registry with ownership and change history. When meaning changes, version or restate deliberately. Do not let a generative catalog description overwrite a governed definition.
AI creates new surfaces: training data, feature stores, prompt context, embeddings, retrieval indexes, evaluation sets, feedback, model output, and human corrections. Data checks must follow each surface.
The voluntary NIST AI Risk Management Framework structures risk work around Govern, Map, Measure, and Manage. NIST states that AI RMF 1.0 is under revision as of 2026; the published 1.0 remains the reference until a successor is issued.
NIST AI 800-4, published in March 2026, identifies categories and open challenges in deployed-AI monitoring, including functionality, operations, human impact, security, compliance, and model behavior. It reports a fragmented, nascent field and open questions; it is not a mandatory standard or complete monitoring method.
For AI pipelines track data coverage, label policy, evaluation integrity, retrieval freshness, unsupported-output rate, harmful feedback loops, and human override. A healthy endpoint latency does not demonstrate a safe or correct answer.
At 07:00, a revenue dashboard shows a 38% decline.
Observability found a technical failure quickly. Finance established semantic and quantitative correctness. The model accelerated investigation but did not approve the repair.
An incident should connect symptom to consumers. Lineage helps identify dashboards, features, regulatory reports, customer exports, and models that may be affected. Classify:
Quarantine or clearly mark suspect assets. Prevent automated consumers from reading a partially repaired table when consistency matters. Record who authorized a backfill and retain before-and-after counts.
Track:
Avoid vanity coverage. A hundred generic null checks on low-value tables do not outweigh one unmonitored regulatory export. Prioritize by decision harm and recoverability.
Data problems often originate in business workflow: an approval skipped, a status entered late, a workaround outside the system. Process mining can reveal where records are created or corrected, but it also depends on trustworthy events. Our process-mining guide explains that reciprocal relationship.
Create a shared map:
This prevents a data team from “fixing” a value that accurately reflects a broken process, or an operations team from blaming a pipeline for inconsistent entry rules.
Before declaring a critical dataset observable, require:
Use our audit-evidence assurance guide to preserve proof that controls ran and exceptions were resolved.
Reviewed 2026-07-30. Main sources are ISO 8000-8:2015, which remains current after confirmation in 2022; ISO/TS 8000-81:2021, confirmed in 2025; the OpenLineage specification overview; the NIST AI RMF page; and NIST AI 800-4. ISO 8000-8 provides concepts and measurement prerequisites, not a universal score. Profiling and lineage reveal structure and movement; they do not alone establish semantic correctness. NIST AI 800-4 catalogs monitoring challenges rather than mandating a complete method.
No. It means selected technical checks passed. Semantic and decision fitness need domain definitions, reconciliation, and use-specific evidence.
Only for narrowly authorized, reversible actions with validated rules and approval proportional to impact. High-impact correction and backfill should remain controlled.
No. Validate coverage. Manual files, BI formulas, notebooks, reverse ETL, and vendor transforms are common blind spots.
Choose one critical metric or model input, map its end-to-end lineage, define its semantics and controls, and rehearse a realistic incident.
Data trust is not the absence of alerts. It is the ability to explain meaning, origin, transformation, uncertainty, consumers, controls, and correction. AI can shorten the distance from symptom to hypothesis. Accountable owners still establish what the data means and whether it is fit for the decision.

AI can help finance and engineering teams detect waste, forecast cloud spend, and connect infrastructure usage to product value.
Read More
Autonomous agents need traces, run histories, approvals, and failure taxonomies so teams can understand what happened after the agent acted.
Read More
A governance-focused guide to using AI for evidence, scenarios, dissent, board briefs, and decision follow-up without outsourcing executive judgment.
Read MoreGet in touch with our team to discuss how we can help your business.