Reading data from a blockchain sounds like it should be simple. The data is public, after all. In practice, pulling a wallet’s full trading history or every transfer of a token is one of the harder problems in Web3 development, and it’s the reason blockchain indexing exists.
This guide explains what indexing is, why raw blockchain data is so awkward to query, how indexing compares to a standard RPC connection, and how to get at historical state without waiting minutes for a single answer. It’s written for people building on-chain, but the first half assumes no prior knowledge.
What Does It Mean to Index a Blockchain?

Indexing is the process of reading raw blockchain data, transforming it, and storing it in a database that’s optimized for search. The result is that queries which would take a node minutes to answer return in milliseconds.
A blockchain stores data as a chain of blocks. Each block holds transactions, and each transaction can emit events (called logs) when a smart contract runs. That structure is great for verifying the chain in order, but terrible for questions like “show me every Uniswap swap this address made in 2024.” To answer that, something has to read every block, decode the relevant events, and organize them so they can be looked up by address, by contract, or by date.
That “something” is an indexer. It watches the chain, processes each new block, and writes the decoded results into a queryable store, usually exposed through an API such as GraphQL or SQL. Developers then query the index instead of the chain itself.
What Is a Subgraph?
A subgraph is an open API that defines how a specific slice of blockchain data should be indexed and served. It specifies which smart contracts to watch, which events to track, and how to map that event data into records you can query.
The term comes from The Graph, a decentralized indexing protocol that popularized the model. You write a subgraph manifest, deploy it, and query the indexed data through a GraphQL endpoint. Uniswap, Aave, and Lido all rely on this pattern to power the data behind their apps.
“Subgraphs are powering some of the most used analytics tools in the space — Uniswap being one of them,” said Hayden Adams, founder of Uniswap, when The Graph launched its network.
Why Is Querying Blockchain Data So Hard?
The short answer: a node is built to validate the chain, not to search it. When you ask a node a question about the past, it often has to do a lot of work to reconstruct the answer, and the standard tools for reading data come with tight limits.
Three problems show up again and again.
Data is stored for verification, not for search. Events are scattered across millions of blocks with no index by contract, token, or user. There’s no “WHERE” clause on a blockchain. To find matching records, you scan.
Reading events doesn’t scale. The main method for pulling historical events over standard RPC is eth_getLogs. It returns the logs in a block range that match a filter, and it’s supported by every EVM node. But the node has to scan every block in the range, so the cost grows with the size of the range, not the number of results. Wide queries time out or get rejected outright.
State older than a few minutes gets thrown away. A standard full node keeps only recent state in fast storage and prunes the rest. Ask it for a wallet balance from two years ago and it either replays transactions from genesis (slow and expensive) or simply can’t answer.
The Limits of eth_getLogs
If you’ve tried to build an event history over RPC, you’ve probably hit a wall with eth_getLogs. The frustrating part is that the wall is in a different place on every provider.
The block-range and result caps are not standardized. One measurement in mid-2026 found the same broad query returned wildly different limits across public endpoints: one capped it at 50 blocks per query, another at 1,000, and a third at 10,000 results. QuickNode limits it to a 10,000-block range on paid plans and just 5 on the free tier, while other providers recommend staying under 5,000 blocks on Ethereum.
So to read a full token history over RPC, you can’t send one request. You have to split the range into windows, page through them, keep a cursor so you can resume, deduplicate on transaction hash and log index, and re-scan the most recent blocks in case of a chain reorganization. That’s a small data pipeline for what feels like a single question, and it’s exactly the work an indexer does for you.
RPC vs Indexing: What’s the Difference?
RPC (Remote Procedure Call) is the direct interface to a blockchain node. It’s how you submit a transaction, check the latest block, or read the current state of a contract. Indexing sits a layer above: it reads data through nodes, processes it, and serves it back through a search-optimized API.
They solve different problems, and most production apps use both.
| RPC access | Indexing | |
|---|---|---|
| What it talks to | A blockchain node directly | A database built from node data |
| Query style | Fixed JSON-RPC methods (eth_call, eth_getLogs) | Flexible GraphQL or SQL queries |
| Best for | Live state, sending transactions, single lookups | History, analytics, filtering and aggregation |
| Historical range | Limited; needs an archive node for old state | Full history, once indexed |
| Cross-contract queries | Not supported | Native |
| Setup effort | Point-and-call | Define what to index, then wait for sync |
A useful rule of thumb: use RPC when you need the present or need to write to the chain, and use an index when you need to search the past or combine data across many contracts. Asking RPC to do an indexer’s job is where latency and rate-limit errors come from.
Who Needs Indexed On-Chain Data?
Anyone whose product answers questions about what has already happened on-chain. If your app only reads current state or sends transactions, plain node access is usually enough. The moment you need history, filtering, or aggregation, indexing earns its place.
The common users:
- Block explorers that show transaction histories, balances, and token holders.
- DeFi dashboards and portfolio trackers that reconstruct a user’s positions and trades over time.
- Analytics platforms studying volumes, liquidity, and on-chain behavior.
- NFT marketplaces tracking ownership, transfers, and sales across collections.
- Compliance and audit tools that need to inspect account state at specific past blocks.
- AI agents, an increasingly large audience — The Graph reported that 37% of new users of its Token API were AI agents querying blockchain data.
How to Access Historical Blockchain Data
Historical blockchain data means the state of the chain at some past point: a balance at block 15,000,000, a contract’s storage last year, an old transaction’s result. There are two practical routes to it, and the right one depends on how much history you need and how often.
What Is an Archive Node?
An archive node stores the complete historical state of a blockchain at every block height, so it can answer a query about any past block instantly instead of recomputing it. It’s the raw-data foundation for historical access.
The trade-off is storage. A full Ethereum node needs roughly 1–2 TB, because it keeps only recent state. An archive node keeps everything, which pushes requirements much higher. Per ethereum.org, archive mode runs around 2 TB on the Erigon client and roughly 12 TB or more on Geth, and it grows every day. Layer 2 networks can be larger still.
| Node type | State retained | Ethereum storage | Good for |
|---|---|---|---|
| Full node | Recent state only (pruned) | ~1–2 TB | Current balances, sending transactions, live monitoring |
| Archive node | Complete history from genesis | ~2 TB (Erigon) to 12+ TB (Geth) | Historical queries, backtesting, analytics, audits |
Running your own archive node gives you unrestricted historical queries but means real hardware, sync time, and maintenance. Many teams instead reach an archive node through a provider that already operates one.
Indexers and Data APIs
An archive node answers historical questions one call at a time, and it still inherits the eth_getLogs scanning problem for events. When you need history that’s already organized, an indexer is the better fit.
Indexing services read from archive nodes, decode the data once, and store it so you can query years of activity with a single request. That’s the difference between fetching a past balance (an archive node handles it well) and pulling a wallet’s entire transaction history filtered by token (an index handles it far better).
In short: use an archive node for point-in-time state at arbitrary blocks, and an indexer for searchable, aggregated history across many blocks or contracts.
How Indexing Protocols Work

Most indexing tools follow the same shape. You describe the data you care about, the service processes matching blocks from genesis (or a chosen start block) forward, and it keeps processing new blocks as they arrive so the index stays current.
The Graph is the best-known example. It uses subgraphs to define what to index and serves the results over GraphQL through a decentralized network of indexers. The protocol has processed over 1.27 trillion queries since launch and supports more than 90 blockchain networks. In December 2025 it shipped a major upgrade, Horizon, that opened its architecture beyond subgraphs to services like streaming data and SQL-based analytics.
It’s no longer the only option. When Alchemy shut down its own subgraph service at the end of 2025, teams moved to a range of faster alternatives, including Goldsky, Envio, and SQD. Some teams also run open-source indexing frameworks like Ponder in-house, or use SQL-first analytics platforms such as Dune for querying rather than app back-ends.
The choice usually comes down to whether you want a hosted API or self-hosted control, how many networks you support, and how fast you need the index to sync.
Accessing Blockchain Data with NOWNodes
Indexing has to read from somewhere, and that somewhere is a node with reliable access to the chain, including its history. That’s the layer NOWNodes provides: node infrastructure across 120+ blockchain networks, so you can read live state and historical data without running the hardware yourself.
For the historical side of the work described above, access to an archive node is what lets you query old balances and contract state or feed an indexer with decoded events from genesis. NOWNodes offers this for major chains, and you can point your indexing setup at an Ethereum endpoint or explore what’s available across the full network list. It’s a way to skip the storage and sync overhead of self-hosting while keeping direct, low-latency access to on-chain data.
The Bottom Line
Blockchain indexing exists because blockchains are built to be verified in order, not searched by topic. Raw node access over RPC is perfect for the present moment and for writing to the chain, but it struggles the instant you ask about the past or want to filter across contracts.
For most applications the answer isn’t one or the other. You use RPC for live state and transactions, an archive node when you need exact historical state, and an indexer when you need that history organized and searchable. Match the tool to the question, and the latency and rate-limit headaches mostly disappear.
FAQ
Is blockchain data free to access?
The data itself is public and can be read from any node, and many providers offer free public endpoints. What you pay for is reliable, high-volume access and the heavy lifting of historical queries or indexing, which need far more resources than a casual read.
Can I query a blockchain with SQL?
Not directly. A blockchain has no query language of its own. SQL becomes possible once the data has been indexed into a database, which is what platforms like Dune and some indexing services provide on top of the raw chain.
Do I always need an archive node for historical data?
Only for exact state at an arbitrary past block, such as a balance at a specific block height. For historical events and transaction history, an indexer built on top of an archive node is usually faster and easier than querying the node yourself.
Can I build an index on top of a public RPC endpoint?
You can start one, but a full historical sync usually strains free public endpoints, which apply the tightest rate limits and eth_getLogs caps. Indexing large histories is smoother against a dedicated endpoint with archive access, since the initial sync makes a heavy, sustained volume of requests.
How long does it take to index a blockchain?
It depends on the chain’s size, the start block, and how much data you’re capturing. Indexing a single contract from a recent block can take minutes, while indexing high-activity contracts across a chain’s full history can take hours or days before the index is caught up to the latest block.



