Every financial agent project starts the same way. Someone says "we'll just scrape it," a week disappears into HTML parsers, and the demo works until the first question that requires a citation. Then the real problems arrive at once: the agent cites a page that has since changed, the index is three weeks stale because nobody built an incremental refresh, and legal asks whether you may keep the embeddings you just spent $400 generating.
Scraping fails for agents for reasons that have nothing to do with difficulty. It fails on provenance — a scraped paragraph has no stable identifier, so "according to the CFO on the Q2 call" is something your model asserts rather than something your pipeline proves. It fails on freshness — a crawler that re-downloads everything nightly is expensive and still lags, while one that only fetches deltas needs a cursor the source rarely provides. And it fails on licensing — the question is not whether you can technically fetch a page but whether you may store a derived representation of it, which is exactly what a vector index is.
Buying data does not automatically solve any of this. Plenty of paid APIs hand you a wall of text with a timestamp and nothing else. The useful question is not "which provider has the most data" but "which provider's output survives contact with a retrieval pipeline." That is what this list ranks.
The five criteria
Every entry below is judged on the same five properties — the ones that determine how much glue code you write.
Chunkability. Does the data arrive in natural units, or as one undifferentiated blob you split with heuristics? A source that hands you semantically coherent pieces has done the hardest part of RAG for you; a source that hands you a 90-page document has handed you a preprocessing project.
Metadata for filtering and citation. Every chunk needs enough attached structure to be pre-filtered before the vector search — so you are not searching 12 million segments for one company's last four quarters — and cited afterwards with a checkable reference.
Incremental sync. Can you ask "what changed since X" and get only the delta? Without a cursor or webhook you are diffing full snapshots, which sets a floor on both cost and staleness.
Latency. How quickly does new information appear after the event? This matters enormously for prices, quite a lot for news, and less than people assume for narrative and fundamentals.
Licensing clarity for derived storage. May you store embeddings, extracted facts and summaries on the plan you are actually paying for? Terms that are silent on this are a risk, not a permission.
1. earningscalls.dev
Best for: narrative data that is already chunked, already cited, and already synced.
This is our product, so treat the ranking with appropriate suspicion — but the reason it sits at number one is a structural property, not a marketing claim. Transcripts arrive pre-split into speaker turns, each tagged with a role (Executives, Analysts, Operator). One turn is one coherent thought by one identified person, which means one embedding chunk: no recursive character splitter, no chunk_size tuning session, no debate about overlap. The discourse structure of the call is the chunking strategy.
Citation metadata comes attached rather than reconstructed — speaker name, role, company, ticker, call date, position in the call — so you pre-filter by ticker, sector, date range or speaker role before touching the vector index, then render a citation a human can verify. Full-text search supports phrase matching, AND/OR and negation, which gives you hybrid retrieval without your own keyword layer. For incremental sync, GET /api/v1/transcripts/recent is cursor-based via since and after_id, so a nightly job ingests only new calls; Enterprise adds webhooks that push them instead. A native MCP server lets an agent query the archive with no tool-calling layer at all.
Coverage is 253,000+ transcripts across 12,799 companies and 175 exchanges, totalling 11.99M speaker-tagged segments, 2020 to today. Pricing is public: free at 50 requests/month, Pro $24.99, Ultra $39.99, Enterprise $299.
Honest limitation: this is earnings call narrative and nothing else. No prices, no fundamentals, no filings, and history starts in 2020. If your agent needs a five-year revenue CAGR or yesterday's close, it needs another source in the stack — which is exactly why the reference architecture below has four boxes and not one. Details in the API docs, and a full walkthrough in feeding transcripts into a RAG pipeline.
2. SEC EDGAR
Best for: authoritative primary documents at zero cost.
EDGAR is the source of record. Every 10-K, 10-Q, 8-K and proxy is there, free, no key required, with full-text search inside filing text back to May 2001. Document-level metadata is genuinely good — CIK, form type, filing date, accession number — giving you durable identifiers that never rot, the single best citation primitive on this list, and incremental sync comes free via the daily and quarterly index files. What it costs you is chunkability: a filing is long, deeply nested and inconsistently formatted, so usable chunks mean exhibit handling, table extraction and section detection tuned per form type.
Honest limitation: the rate is capped at 10 requests per second per user with a declared user-agent, full-text results are not pageable past 10,000 hits, and the index does not reach pre-2001 filings even though EDGAR holds older ones. On preprocessing burden it is the hardest source here — but on authority and price it beats everything, including us.
3. Financial Modeling Prep
Best for: structured fundamentals across a wide universe.
If your agent needs to answer "what was gross margin in each of the last twelve quarters," FMP beats anything narrative. Income statements, balance sheets, cash flow, ratios and key metrics arrive as clean JSON with long histories for large caps across a broad global universe. Chunkability is the wrong frame here; what matters is that each record is already a row keyed by period and ticker, so it joins cleanly to everything else in your warehouse. Bulk endpoints on the higher plans pull whole datasets instead of looping per ticker.
Honest limitation: coverage quality is strongest for US and European equities and thins out for smaller or less-liquid listings, so verify the names you care about first. Bulk access sits on the upper tiers. For structured fundamentals this beats us outright — we do not offer them at all.
4. Polygon.io
Best for: market data with developer ergonomics that respect your time.
Polygon is what a market data API looks like when someone has actually built on one: consistent response shapes, sane pagination, WebSockets for streaming, and flat files over an S3-compatible endpoint for bulk historical loads, so you are not making a hundred thousand REST calls to backfill. That flat-file path is also the incremental-sync story — pull the daily file, load it, done.
Honest limitation: it is market data, full stop — no narrative, no filings, and fundamentals are not the reason to be here. Real-time access and the deeper historical asset classes sit on paid tiers; check which classes your plan includes rather than assuming.
5. Alpha Vantage
Best for: prototypes and macro context on a small budget.
Alpha Vantage spans a wide surface — equities, FX, crypto, commodities and US economic indicators drawn from Federal Reserve and BLS series — behind one simple key-based API. For an agent that needs to say something sensible about the rate environment alongside a company view, having macro series in the same client as prices is a real convenience.
Honest limitation: the free tier is a demo, not a foundation — 25 requests per day at 5 per minute will not sustain a pipeline. Premium tiers start at $49.99/month for 75 requests per minute and drop the daily cap. Per-call rate limits shape your ingestion design more than the data does, so plan for a queue.
6. A news API (Benzinga, NewsAPI, or similar)
Best for: the freshness layer, if and only if the licensing works for you.
News is the lowest-latency input an agent can have and the one most likely to change an answer. Headlines are naturally chunkable — a headline plus summary is about the right size for an embedding, with a timestamp and usually a ticker attached — and incremental sync is easy, since the natural query is "everything since this timestamp."
Honest limitation: licensing is the whole story here and varies sharply by provider and plan. Some terms permit caching only briefly, some restrict how much article body you may retain, some price storage and display rights separately from API access. A vector index of article text is a derived copy, and whether you may keep one is a question to settle before you build. This is the one category where "we'll sort it out later" reliably turns into rework.
7. AlphaSense or Quartr
Best for: breadth and depth that nothing else on this list matches, at professional prices.
These platforms aggregate filings, transcripts, broker research and expert content with excellent search on top, and for a human analyst they are outstanding. AlphaSense offers a developer API suite with official SDKs, so programmatic access is real rather than theoretical. If your agent serves a team that already holds seats, integrating there is often the shortest path to the widest corpus.
Honest limitation: neither publishes list pricing; Pro and API access are sales-negotiated seat-based contracts, and the content licensing that makes them valuable is also what constrains what you may store downstream. They rank seventh purely on this post's criteria — accessibility and predictability for an agent builder — not on data quality, where they are at or near the top.
Reference architecture
A realistic stack for a financial research agent uses four sources, joined on two keys.
- Prices and volumes from a market data API, streamed or loaded as daily flat files.
- Fundamentals from a fundamentals API, refreshed after each reporting season.
- Filings from EDGAR, chunked once and re-processed only when a new accession appears.
- Narrative from earningscalls.dev, pulled incrementally via the
recentcursor or pushed by webhook.
Everything joins on ticker and date. That is the entire integration design. "Margin compressed in Q2 — what did management say about it?" becomes a fundamentals lookup that identifies the quarter, a filtered retrieval over transcript segments for that ticker and call, then an answer whose every claim carries a speaker and a date.
Pull segments directly to see the shape of a chunk:
curl "https://earningscalls.dev/api/v1/speakers/257198?speaker_type=executive" \
-H "X-API-Key: $KEY"
Each returned segment is a ready-made chunk: the text is one speaker's turn, and the citation metadata — name, role, company, ticker, date, position in the call — is already attached, so you embed the text and store the rest as payload with no extraction step. The same pattern in agent form is covered in building an earnings research agent with MCP and Claude.
Keep a source, source_id and retrieved_at on every record in every store. When an answer is wrong six months from now, that triple tells you which pipeline produced the bad chunk.
On licensing derived storage
One paragraph, because it matters more than its length suggests. Storing derived records — embeddings, extracted metrics, summaries, an index built from data you pay for — is normal and permitted on paid plans at most providers, including ours. Redistributing the archive itself is not: reselling the corpus, exposing a bulk download to your users, or shipping a product whose value is the raw data rather than what you built on it. The line sits between using the data inside your application and republishing it, and every provider on this list draws it slightly differently. Read each set of terms before you design your storage layer rather than after — retrofitting a retention policy into a live vector store is unpleasant work. Ours are at /terms.
Related comparisons
- Best MCP Servers for Investors and Financial Research — eight MCP servers, and how to combine prices, fundamentals and narrative
- Best Earnings Call Transcript APIs for Developers — eight transcript APIs ranked on coverage, structure and integration path
If earnings call narrative is the layer your agent is missing, start at earningscalls.dev — the free tier is enough to test the chunking properties before you commit to anything.