WHAT YOU’LL TAKE AWAY
- Make data improvement part of a valuable operational workflow.
- Keep observed facts, proposed matches, and inferred values distinguishable.
- Validate, approve, record, and recheck corrections before relying on them.
Utilities have valuable information scattered across geographic information systems, work orders, spreadsheets, scanned drawings, meter records, and people’s notes. Inconsistent identifiers and missing fields make that information expensive to use. They do not make it worthless.
AI can help with the cleanup itself: finding likely duplicates, extracting attributes from documents, proposing field mappings, and prioritizing inconsistencies for review. This changes the starting point. Data improvement can happen alongside useful work, with progress measured in records made usable and decisions supported.
The critical distinction is between recovering evidence and inventing certainty. A missing value remains unknown until there is an acceptable basis for filling it.
Start with a problem that matters
Choose a bounded area where poor data delays a real task. Examples include inspections that cannot be matched to assets, duplicate service requests, or equipment records missing an attribute needed for planning. Define the quality needed for that task.
A planning screen may tolerate an explicitly labeled estimate. A safety-related operating decision may require verified information. “Complete” and “fit for use” are different tests.
Baseline the current burden: unmatched records, duplicate rates, staff time spent searching, and the number of decisions blocked. This ties cleanup to an outcome that an operational owner can support.
Use the right technique for each problem
| Data problem | How AI can help | What establishes trust |
|---|---|---|
| Different names for one asset | Propose matches using identifiers, location, and context | Reviewed matches and checks for false merges |
| Attributes buried in documents | Extract candidate values with page references | Original evidence and field-level review |
| Inconsistent labels or units | Suggest a common mapping | Approved dictionaries and deterministic conversion |
| Conflicting records | Explain the conflict and gather supporting evidence | A field owner and an explicit authority rule |
| Missing or implausible values | Detect gaps and prioritize investigation | Verification, inspection, or a clearly labeled estimate |
Use straightforward rules wherever they are sufficient. AI adds the most value when context, unstructured material, or ambiguous relationships make those rules difficult to write.
Follow an evidence-preserving correction loop
Identify. Scan for duplicates, missing fields, impossible combinations, stale records, and broken relationships. Group related issues so reviewers can fix a cause rather than handle every symptom separately.
Attribute. Show the source object, current value, capture date, and associated evidence. Keep the original record intact. A proposed link between systems should be distinguishable from a confirmed identity.
Propose. Present the correction and its rationale. Separate an observed value extracted from a legible nameplate from an inferred value based on similar assets. A model’s confidence is an input to review, not proof of correctness.
Validate. Check format, units, allowed values, dependencies, topology, and the business rules relevant to the workflow. Use engineering tools where physical consistency must be checked. Passing a schema check alone cannot establish that an asset attribute is true.
Approve and apply. Route consequential or ambiguous changes to the appropriate person. For well-understood low-consequence corrections, a utility may authorize automation within narrow rules. Record what changed, why, and under whose authority.
Verify and learn. Confirm that the destination accepted the change. Revalidate affected records, retain correction history, and route failed updates or unresolved issues to an owner. Feed confirmed decisions back into matching and validation rules.
This loop reflects the workbook’s data management sequence: issue detection, source attribution, correction validation, accepted updates, error handling, history, and human follow-up.
An illustrative asset example
A transformer appears under different identifiers in GIS and a maintenance spreadsheet. A scanned inspection includes a serial number and a location. AI proposes a link among the three records and highlights that their ratings disagree.
The useful result is an evidence package: candidate identity, matching reasons, conflicting values, and the relevant inspection image. An engineer confirms the identity, determines which rating is supported, and approves a correction. The system updates the permitted destination and checks the result.
If the serial number is unreadable, the process creates a field verification task. It does not fill the blank simply to make the dataset look complete.
Measure quality as a continuing operation
Track false matches, accepted corrections, unresolved exceptions, repeat defects, and review effort. Sample accepted changes to catch errors that a reviewer or rule missed. Monitor whether new imports reintroduce the same problem.
Start with one region, asset class, or document set. Once the correction loop works, expand it. A useful foundation keeps improving as new data arrives—and every improvement should remain explainable.
Sources & further reading
Practical guidance combines the source material below with editorial analysis. Examples and suggested approaches are illustrative.
- Senpilot website: AI-powered data harmonization
The provided website illustrates matching assets across GIS, OMS, SCADA, and AMI and flagging missing attributes.
- Senpilot Global List of AI Use Cases in Utilities, September 2026
Draws on DM01–DM16 data management workflows and E25 data quality management.
