[data]

Rebuilding a Legacy BI Estate on Microsoft Fabric with a Multi-Agent AI Pipeline

How an enterprise analytics organization moved a hand-migrated reporting estate onto a governed Microsoft Fabric foundation — and left its own team running the practice.

How an enterprise analytics organization moved a hand-migrated reporting estate onto a governed Microsoft Fabric foundation — and left its own team running the practice.

The challenge

The client’s analytics organization maintained a large legacy reporting estate that was being migrated one report at a time, by hand, by its most specialized engineers.

The work itself was high quality. The delivery model around it was the problem.

Reports could be downloaded, modified, and republished with no shared review practice. Architecture, ownership, model quality, and release readiness varied across the estate. Business meaning lived in individual reports rather than in a governed model, which meant that any AI experience built on top of the estate would interpret the same question differently from the next one.

Manual migration compounded all of it. It consumed scarce specialist capacity, repeated the same decisions report after report, and allowed mappings to drift as each report was interpreted in isolation.

The organization needed migration to scale. It also needed the result to be governed well enough to build AI on.

The approach

A legacy report cannot be migrated by copying its charts. Its meaning sits in queries, joins, calculated columns, measures, filters, parameters, and assumptions that were never written down. The target environment is architecturally different, so every source concept has to be related to the correct entity, measure, and business definition.

Pointing a language model at the estate and asking for a rewrite produces output that looks correct and fails quietly. We took a different route: build the foundation first, then rebuild migration as an engineering pipeline where each stage can be evaluated, rerun, and improved on its own.

The principle held across all three workstreams. Deterministic software does what is mechanical. AI does what requires interpretation. People decide what is ambiguous or consequential.

Workstream 01 — A governed engineering system

We moved analytics assets out of opaque binary files and into source-controlled code, then built the delivery practice around them: Git branching, pull requests, named ownership, CI/CD, automated quality gates, parity validation, owner sign-off, and rollback.

We enriched the semantic models so AI could use them reliably, and transferred the practice to the client’s engineers through working guides, office hours, and reusable automation.

This workstream came first by design. Agent quality is capped by the foundation underneath it.

Workstream 02 — AI-accelerated report migration

A staged pipeline that moves legacy reports into the governed environment with evidence attached at every step. Six stages, each independently evaluated.

01 — Extract. Deterministic code decomposes each legacy report into structured metadata: sources, queries, fields, calculations, filters, visuals, layout. No model is involved. This becomes the factual record everything downstream trusts.

02 — Interpret. An agent establishes what the report was actually for — business purpose, the measures and dimensions presented, the role of each filter, the meaning of abbreviated field names. The result is kept as an explained recommendation tied to evidence.

03 — Map. An agent maps legacy concepts to the governed semantic model, weighing names, descriptions, types, expressions, and business terminology. Where evidence is thin, it surfaces an explicit gap rather than forcing a plausible match.

04 — Generate. The system produces the target report as an inspectable, source-controlled artifact against the governed model — pages, visuals, approved measures, filters, formatting — with migration decisions and unresolved items attached.

05 — Validate. Deterministic tests cover mechanical facts, AI handles semantic comparison, and source-to-target reconciliation checks expected totals. A report that looks right but uses the wrong measure fails inside the pipeline rather than in production.

06 — Route. Anything ambiguous or consequential reaches a specialist with the source evidence, the interpretation, the candidate mappings, and the validation findings already assembled.

Workstream 03 — Reusable semantic intelligence

Understanding the canonical model well enough to migrate reports produced knowledge other teams wanted for their own AI experiences. Rather than have each team rebuild the catalog and drift out of sync, we centralized that context and exposed it as a governed service.

A person, an application, or another agent submits a plain-language request. The service turns it into an inspectable analytical plan, validates that the combination is supported, executes against governed data, and returns a defined contract: interpreted intent, selected model objects, execution plan, answer, result data, visual specification, field mappings, and validation findings.

Answers are anchored to approved measures. The language model interprets and explains. It is never the source of analytical facts.

The architecture

Four layers, with a trust plane running across all of them: identity, permissions, evaluations, tracing, and human gates.

Experience: Fabric app with SSO, Teams and assistants, API consumers.

Agent and orchestration: Orchestrator, specialized agents, tool contracts (MCP), structured outputs.

Semantic and context: Canonical semantic model, terminology and rules, constraints and defaults.

Data platform: Microsoft Fabric, source systems, deterministic services.

No request goes straight to execution. Every request becomes an inspectable plan first, so it can be validated, logged, and replayed.

Results

All three tracked report waves were converted to source-controlled formats, creating a reviewable and recoverable estate.

Ten report-level releases were recorded in Azure DevOps, each with release method, approval, rollback path, and post-release proof.

105 pull requests were raised across four contributors. In one later sprint, 83% of pull requests came from the client’s own team rather than our embedded engineers.

A governed data agent was published and validated in Fabric and Teams, alongside an application shell deployed with single sign-on behind explicit production-readiness gates.

The last figure is the one that mattered most to both sides. The engagement was scoped to leave a team that ships without us.

Why the accuracy holds

Accuracy did not come from assuming AI is more reliable than people. It came from controls that are difficult to apply consistently by hand.

Source facts are extracted deterministically and become the single factual record for the whole migration.

Governed measures are reused rather than recreated as report-local logic.

Low-confidence and unsupported mappings are surfaced, not quietly resolved.

Every generated element traces back through mapping, target object, and validation result to its source definition.

Human approval stays mandatory wherever business meaning is ambiguous or consequential.

The outcome

What began as a migration became a department-scale platform pattern: governed data, repeatable delivery, AI-ready semantic models, controlled AI experiences, and a client team operating the practice independently.

Specialist time moved off routine reconstruction and onto exception judgment. Mappings became consistent and reusable. Quality checks moved earlier, so rework was caught before it compounded. Downstream users received modernized reports that preserved the analytical meaning their decisions depended on.

The pattern applies well beyond this estate. Any decade-old platform replacement is a semantics migration before it is a software project — the embedded business rules are the real problem, and extracting their meaning is the work.

Frequently asked questions

Can AI migrate legacy BI reports accurately?

Not on its own, and not in a single pass. Accuracy comes from the architecture around the model: deterministic extraction of source facts, AI used for interpretation and mapping, validation and reconciliation inside the pipeline, and a human gate on anything ambiguous. A single-prompt rewrite produces output that looks correct and fails silently.

What is a semantic layer, and why does AI need one?

A semantic layer is the governed definition of what a business’s measures actually mean — entities, relationships, approved measures, aggregation behavior, valid analytical paths, terminology. Without one, each AI experience invents its own interpretation and they drift apart. With one, every answer can be anchored to an approved definition and inspected afterward.

What makes a multi-agent system different from a chatbot?

A chatbot answers questions. A multi-agent system coordinates specialized agents performing work against governed systems, with an orchestration layer, inspectable plans, validation, and approval gates. Different problem, different architecture.

How long does a Microsoft Fabric migration take?

It depends on estate size and how much business logic is undocumented. The more useful question is what gets built first. In this engagement the governed engineering system — source control, quality gates, validation, rollback — was sequenced before any AI-assisted authoring, because agent quality is capped by the foundation underneath it.

Talk to us about your estate

We work with organizations replacing legacy analytics platforms, consolidating fragmented reporting estates, and moving stalled AI pilots into production.

Book a 30-minute architecture review →

Delivered for an enterprise data and analytics organization. Client identity withheld under confidentiality.

Most teams stop at the plan. This one didn’t.

Most teams stop at the plan. This one didn’t.

Most teams stop at the plan. This one didn’t.

Let's make AI [real] together.

Let's make AI [real] together.

Let's make AI [real] together.

[up]

start.13

lift

grade

level

scale

skill

focus

start.13