New: AI-ready data insights featuring BlackRock. Read more
Solutions / Data Intelligence

Data Intelligence

Turn statements into structured data, including the custodians that publish no feed.

0:00 / 0:00
The problem

A feed is never guaranteed. A statement always is.

Every automated data flow assumes the custodian publishes a feed. A feed is a commercial and technical decision somebody else makes, and for a meaningful part of any multi-custody book that decision has gone the other way.

When a feed exists

An API, SWIFT, or FIX feed from the custodian. Positions, transactions, and valuations arrive on a schedule, reconciled, and nobody touches them. This is what WealthArc Data Feeds already does across a network of more than 170 institutions.

When there is no feed

For any of these reasons a live feed may simply not exist, and waiting for one is not a plan.

  • A new relationship, still mid-onboarding
  • A smaller custodian with no published interface
  • A jurisdiction where feeds are not the norm
  • Held-away and legacy holdings
  • Anything that predates the connection

What you can always get is a statement. WealthArc Data Intelligence reads that statement and produces the same structured data a feed would have delivered, so a custodian that will not connect stops being a hole in the book.

The product

WealthArc Data Intelligence.

WADI is the first of WealthArc's released AI agents. It reads financial documents, extracts and standardizes every field, validates the result against the reference data already in your WealthArc Data Box, and delivers structured output to whatever runs downstream. It is a module of the infrastructure, not a tool beside it.

WealthArc
Data Intelligence
Abbreviated WADI. The first of WealthArc's released AI agents, and the one that closes the last mile of wealth data processing.
Sits on
Connector and WealthArc Data Box, the connectivity and storage layers you already run on.
Delivered through
WealthArc Portfolio Management and Portfolio Viewer, or Prism, the API and MCP distribution layer.
Reads
PDF documents. More document types are being added.
What it does

One statement in. Positions, transactions, and valuations out.

The engine classifies the document, finds the regions that carry data, extracts every field into a predefined schema, and checks the result against your own portfolio and reference data before anything is released.

Source documentPDF · 4 pages
Portfolio statementCustodian, 31 August 2026
Received via SFTP
01
Account
4471‑02‑118
Portfolio
Duarte Family Trust
Valuation date
31.08.2026
Currency
USD
02
HoldingISINQuantityMarket value
Global Equity FundXX000000010112,4003,118,880
Corporate Bond 2.25% 2031XX00000002081,500,0001,442,250
Money Market FundXX00000003159,880988,120
Emerging Markets EquityXX00000004224,2601,704,000
and 30 further positions
03
Cash balance
1,159,410 USD
Accrued interest
8,204 USD
Total portfolio valueUSD 8,412,660
01  Portfolio object4 fields
"accountNumber": "4471-02-118", "portfolioName": "Duarte Family Trust", "valuationDate": "2026-08-31", "referenceCurrency": "USD"
02  Position objects34 rows
"isin": "XX0000000101", "quantity": 12400, "price": 251.52, "marketValue": 3118880.00, "assetClass": "Equity fund"
03  Cash and validation3 checks
"cashBalance": 1159410.00, "accruedInterest": 8204.00, "positionsTotal": 8412660.00
Totals match the statement ISINs resolved Prices in range
WISDOM objects Into WealthArc Data Box Full audit trail

DisclaimerThe interface shown is an illustration, not a live product. All names and figures shown are sample data. ISINs are placeholders and do not identify real securities.

The pipeline

Six stages, none of which you operate.

A document arrives and comes out the other end as data in your WealthArc Data Box, on the same path every custodian feed takes. There is no configuration to write and no template to maintain per custodian.

Document to data

The same route a feed takes, entered one stage later.

Document PDF, in whatever shape the sender's system produces. Arrives
AI parsing Classified by document type, then read field by field against a predefined schema. WADI
Structured data Clean JSON: identifiers, dates, metrics, and every table row. JSON
Quality checks Presence, plausibility, and reconciliation against your reference data. Anomalies flagged. Validated
ETL transformation Mapped into WISDOM investment objects, the same canonical shape every feed produces. WISDOM
WealthArc Data Box Stored with full lineage, ready to be read by anything you authorize. Landed
WealthArc Portfolio Management WealthArc Portfolio Viewer Any PMS Core banking Reporting platforms Client portals Data warehouses AI agents

WealthArc clients read it directly in Portfolio Management and Portfolio Viewer. Everything else is served through Prism, the API and MCP distribution layer, so every system you run reads the same validated data rather than its own copy.

Getting documents in

Three ways in, two of which nobody has to remember.

Once a channel is configured it runs in the background, so documents arrive without anyone sending them. Manual upload stays available for the one-off that turns up by email.

A  Automated collection
Pulled from where they already sit

WealthArc Data Intelligence collects documents from your own tenant-side storage on a schedule, so the operations team keeps working the way it already works. Nothing has to be forwarded.

SFTP Microsoft OneDrive Dropbox
B  Direct from custodians
Pushed to your secure store

Custodians deliver documents straight into dedicated, secure storage inside your WealthArc environment, which is the closest thing to a feed a custodian without one can offer.

Custodian push Dedicated storage
C  Manual upload
Dragged in, when it has to be

Drag and drop, either in the Data Intelligence interface or embedded inside WealthArc Portfolio Management, for the document that arrives by email the day before a committee meeting.

Standalone Inside WealthArc Portfolio Management
What comes out

Fields, not a summary of the document.

Extraction is against a predefined schema per document type, covering the identifiers and dates that place the document, the headline metrics, and every row of every table it contains.

IdentificationWhat the document is, and whose it is
documentTypecustodianaccountNumberportfolioName portfolioReferenceclientReferencereferenceCurrency
DatesThe period the data belongs to
valuationDatestatementDateperiodStartperiodEnd bookingDatevalueDate
Headline metricsThe figures a report leads on
totalPortfolioValueassetsUnderManagementnetContribution performanceTWRperformanceMWRaccruedInterest
PositionsEvery row of the holdings table
isininstrumentNameassetClassquantity pricepriceCurrencymarketValuemarketValueRefCcy costPriceunrealisedPnLweight
TransactionsEvery row of the movement table
transactionTypeinstrumentquantityprice grossAmountfeestaxesnetAmountcounterparty
Cash and flowsBalances, movements, and FX
cashBalancecurrencycashMovementTypeamount fxRateincomeType

NoteField names are illustrative of the schema shape. The exact schema per document type is set during implementation, and the output maps into WISDOM investment objects either way.

Document coverage

Custodian documents, and the illiquid side too.

Custodian and portfolio documents
The bulk of a multi-custody book

The document types that carry most of the data in a multi-custody book, delivered as PDF.

Custodian account statements Portfolio and investment reports Custody and valuation reports Performance reports Transaction and cash movement statements
Illiquid and private markets
The side that never had a feed

Private markets paperwork arrives as PDF almost everywhere, and is still read by hand in most places. Parsed into structured positions, the illiquid side of the balance sheet lands in the same record as everything else.

Capital call notices Distribution notices Commitment schedules NAV statements Cashflow reports
Quality and verification

Extraction you can put your name to.

Reading a document is the easy half. The half that decides whether the output is usable is what happens next: three layers of checking, and an optional human sign-off before anything reaches a downstream system.

Layer 1
Presence

Is every field the schema requires actually there? A statement missing a valuation date or a reference currency is caught before it is parsed further, not after it has been loaded.

Layer 2
Plausibility

Are the values possible? Prices outside a sensible range for the date, quantities that do not match a known instrument, and totals that do not add up are flagged rather than passed on.

Layer 3
Reconciliation

Does it agree with what you already hold? Extracted values are checked against the reference and portfolio data in your WealthArc Data Box, which is the check a standalone tool cannot make.

Verification queue31 August 2026 · 5 documents
Custody statementAccount 4471-02-118
34 positions
0 flags
Released
Portfolio report, Q2 2026Account 4471-02-118
18 metrics
0 flags
Released
Valuation reportAccount 5510-88-004
12 positions
0 flags
Released
Transaction statement, AugustAccount 5510-88-004
61 movements
0 flags
Released
Performance report, H1 2026Account 4471-02-118
9 metrics
1 flag
Held for review

Flagged at layer 2: a price outside the plausible range for the valuation date. Nothing on this document is released until it is resolved, and the resolution is recorded against the field.

Audit trail Every field, every correction, and every approval
Sign-off Optional, and configurable per document type

DisclaimerThe interface shown is an illustration, not a live product. All names and figures shown are sample data.

AI on your own data

Better data in, better AI out.

An agent is only as good as what it reads. Point one at a folder of PDFs and it will answer confidently and inconsistently, because there is nothing underneath the answer. Normalization and validation are not housekeeping steps ahead of the interesting part; they are what makes the interesting part trustworthy.

Without structured data
Raw PDFs and fragmented feedsLLM or agent Unreliable output
With WealthArc and WADI
Structured, normalized, validated data LLM or agent, through Prism Answers you can act on

The same holds for your own models. Data stays inside your WealthArc environment and is readable only by the models and agents you authorize, over the API or the MCP server.

Why not a general purpose reader

An off-the-shelf model has nothing to check its answer against.

Any modern model will read a statement and return something plausible. The question is what happens when it is wrong, and whether you would know.

Off-the-shelf model
WealthArc Data Intelligence
Document taxonomy
No wealth management document taxonomy. Every custodian layout is a new problem, solved with a prompt.
Purpose-trained on wealth management document types, so a custody statement is recognised as one.
Validation
Nothing to validate against. A confident wrong number looks exactly like a correct one.
Validated against your live portfolio and reference data in WealthArc Data Box, and flagged when it disagrees.
Integration
A standalone tool, with its own integration to build, secure, and maintain.
A native module of the infrastructure you already run on. No new integration project.
Consistency
Output quality varies by document type and format, and the variation is hard to see at volume.
One consistent output whatever the source document, mapped into WISDOM investment objects.
Delivery to AI
Whatever you build around it.
Prism, a modern API and MCP layer built for agent consumption from the start.
Historic data

Years of statements are years of analysable history.

Almost nobody delivers this well, which is why almost everybody has the problem. Because the same engine can parse a statement from 2015 as easily as one from last month, history stops being a box of PDFs and becomes data in your WealthArc Data Box, at the same quality as anything arriving today.

One portfolio, before and after a live feed
’15’16’17 ’18’19’20 ’21’22’23 ’24’25 ’26
← Eleven years reconstructed from statements Live feed from 2026 →

DisclaimerThe chart shown is an illustration, not a live product. All names and figures shown are sample data.

Sales enablement
Win the pitch in the same meeting

An asset manager wants a new client. Instead of reading years of statements by hand, they parse the prospect's existing statements and get structured historic data straight away: enough to run real analytics and build a data-driven proposal on the spot. What took days becomes a same-meeting turnaround.

Regulatory and continuity
Reporting from before the relationship

Clients need historic performance and holdings that predate their relationship with their current provider, either to meet record-keeping requirements or to give a new client continuous reporting that includes the history from before they onboarded. Today that data is trapped in old statements. This makes it usable again.

Data portability
No lock-in by default

Because WealthArc Data Box can absorb historic data regardless of source, nobody is stuck with a data provider purely because that provider is the only one holding their history. That removes a switching barrier which most providers create simply by being unable to backfill.

“Your current provider is the reason you feel locked in. We remove that.”

Built for

Built for how wealth data actually moves.

Private banks and EAMs

Digitize custodian statements across a multi-custody book without re-keying, and close the accounts where no feed will ever exist.

Family offices

Turn capital calls, distributions, and NAV letters into structured positions, so the illiquid side of the balance sheet stops living in a spreadsheet.

Asset managers and fund administrators

Automate extraction for regulatory and investor reporting, for managers and administrators alike, on the document types that arrive as PDF by default.

Multi-custody wealth managers

One consolidated data layer regardless of which custodian produced the statement, so consolidation is not limited to the custodians that happen to publish an interface.

Trustees and trust or foundation service companies

Consolidate trust and foundation holdings across custodians into one clean, auditable record, with the supporting statements and deeds stored alongside the data they produced.

WealthTech and FinTech platforms

Take structured output through Prism and ship a product on top of it, rather than building document parsing and a validation layer of your own.

Getting started

The last mile should not be the last hurdle.

Already a WealthArc client
A module, not a project

WealthArc Data Intelligence plugs in on top of your existing Connector and WealthArc Data Box setup. There is no new integration to build, no new store to secure, and the output lands in the same place your feeds already land.

New to WealthArc
Start on one document type

Begin with a scoped pilot on your highest-volume document type, with commercial terms to match. If the accounts with no feed are the pain, start there; if it is eleven years of statements, start there instead.

Use cases

What people build with it.

Custodians That Publish No Feed

Close the gap where a custodian offers no API, SWIFT, or FIX feed at all, and never will. WealthArc Data Intelligence ingests the PDF statements instead and delivers them as a document based data feed, the same structured positions, transactions, and valuations a live feed would have produced, so held-away portfolios and smaller or newly onboarded custodians stop being keyed in by hand.

Historic Data Digitization (PDF → Structured Data)

Parse years of legacy custodian bank statements into clean, structured historic data with WealthArc Data Intelligence, instead of a manual read-through. Turn a prospect pitch into a same-meeting, data-driven proposal, satisfy regulatory record-keeping requirements, and let clients switch providers without losing their history.

Trust Portfolio Oversight, Document Vault & Reporting

Regularly check the underlying portfolios held within every trust, keep the supporting statements, deeds, and legal documents safely stored alongside the reconciled data, and generate beneficiary and regulatory reports on demand - one trusted record per trust, not a filing cabinet.

Other solutions

Need something else?

Try Data Intelligence

Send us one statement from a custodian that will not connect, and we will send back the structured data.

Get started →