Onchain vs offchain data: what’s the difference?

Onchain data is information stored directly on a blockchain — every transaction, wallet balance, and smart-contract state that the whole network keeps and can verify. Offchain data is everything kept somewhere else: a company’s servers, a file on IPFS, a price feed, or a Layer-2 batch that hasn’t settled yet.

The short version is a trade-off. Onchain data is expensive to store but trustless and permanent. Offchain data is cheap and flexible, but you have to trust wherever it lives. Almost every real application uses both, and the skill is knowing where to draw the line.

That line isn’t always obvious, and it has moved over time as blockchains got better at handling data cheaply. This guide covers what each type is, why the split exists, who relies on it, and how data crosses from one side to the other.

What is onchain data?

Onchain data is any information written into the blockchain itself and copied across the network. Once a block is confirmed, that data sits on thousands of machines at once, and anyone can check it independently without asking permission.

A few things count as onchain data:

  • transactions (sender, recipient, amount, timestamp);
  • account balances and smart-contract state variables;
  • event logs emitted by contracts when something happens;
  • block metadata like timestamps, gas used, and validator signatures.

Take a simple ERC-20 token transfer. The blockchain records who sent the tokens, who received them, how many moved, and which block it landed in. It also updates the token contract’s internal ledger so the new balances are the ones everyone sees next. All of that is onchain — replicated, timestamped, and open to inspection.

Here’s the property that makes it valuable: nobody has to take your word for it. If a transfer is onchain, anyone helping run the network can confirm it happened, and no single party can quietly change it later. That verifiability is the reason to put something on a blockchain — and, as we’ll see, the reason it costs so much.

What is offchain data?

Offchain data is everything a blockchain application uses that doesn’t live on the chain. The blockchain might point to it, but the data itself sits elsewhere — and “elsewhere” can mean very different things.

Common examples:

  • the image and metadata behind an NFT, usually on IPFS or Arweave;
  • price feeds, weather data, or sports scores supplied by an oracle;
  • KYC records, user profiles, and application databases;
  • historical analytics and indexed query results;
  • a rollup’s transaction data before it settles to Ethereum.

Some offchain data lives on ordinary centralized servers — fast and cheap, but only as reliable as the company running them. Other data goes to decentralized storage. On IPFS, files are addressed by a content identifier (CID), a hash of the file itself, so the same content always resolves to the same address. The catch is that IPFS keeps a file available only while someone “pins” it; stop paying for pinning and the file can disappear. Arweave takes a different route, charging once to store data with the goal of keeping it available for the long term.

Often the chain stores only a cryptographic hash — a short fingerprint — of the offchain data. The hash proves the data hasn’t changed, while the heavy content stays cheap to host off the chain. That hash-and-pointer pattern shows up everywhere once you start looking for it.

Why isn’t everything stored on the blockchain?

Because onchain storage is deliberately scarce and expensive. Every byte you write has to be stored and replicated by the entire network, effectively forever, so the protocol prices it high to stop the chain from bloating.

The numbers are blunt. Writing a single 32-byte slot to Ethereum storage costs about 20,000 gas (ethereum.org). At around 15 gwei, with ether in the low-thousands-of-dollars range, that one slot costs roughly a dollar — and a single kilobyte, about 32 slots, lands in the tens of dollars. Store a photo that way and you’re paying hundreds or thousands of dollars for something an ordinary server would host for a fraction of a cent.

Not all onchain data costs the same, though. Permanent storage is the priciest tier. Transaction calldata — data attached to a transaction but not saved in contract storage — runs about 16 gas per non-zero byte, which works out roughly forty times cheaper per byte than storage. Blob data, introduced for rollups, is cheaper still. That cost ladder is exactly why developers think hard about where each byte goes.

There’s a second reason beyond your bill. Every participant carries that data too, so unchecked growth makes the chain heavier for everyone who helps run it. That’s why Ethereum’s designers keep looking for ways to shrink what the network must hold. As Vitalik Buterin put it in a 2021 Reddit AMA, “Statelessness + PBS would allow independent validators to run with basically no storage requirements. Only builders and light client servers would have storage requirements.” When the people building the protocol work this hard to store less, that tells you how real the cost is.

So developers follow a simple rule. Keep only what must be trustless and permanent onchain, and push everything large, private, or fast-changing offchain — anchored by a hash when integrity matters.

How onchain and offchain data compare

The choice usually comes down to a handful of properties. This is how the two stack up:

What mattersOnchain dataOffchain data
Where it livesOn every machine in the networkServers, IPFS/Arweave, oracles, Layer 2s
Cost to storeHigh — priced in gasLow, often near-free
PermanenceEffectively permanent and immutableLasts as long as someone hosts it
VerifiabilityAnyone can verify it independentlyYou trust the source (or a hash of it)
PrivacyPublic by defaultCan be kept private
Who controls itThe network as a wholeWhoever hosts or serves it
SpeedLimited by block timesAs fast as ordinary web infrastructure
Best forOwnership, balances, settlementMedia, documents, feeds, analytics

The pattern in that table is the real lesson: you’re trading trust for cost. Onchain, you pay a lot and get guarantees nobody can override. Offchain, you pay almost nothing and accept that someone — a company, a pinning service, an oracle — has to stay honest and online. Neither is “better.” They answer different questions.

Most production systems don’t pick a side; they combine both in a hybrid design. The trust-critical facts go onchain, the bulky or private content goes offchain, and a hash links the two so tampering is detectable. Get that division right and the weaknesses of each side mostly cancel out.

Who uses each type of data?

Almost every serious blockchain product mixes the two. The interesting part is where each one draws the line.

Wallets and explorers

Wallets and block explorers are mostly onchain readers. They pull balances, transaction history, and contract state straight from the chain and show it to you. Reading that data at scale is its own challenge, which is why teams lean on hosted infrastructure — NOWNodes’ block explorers and indexed data APIs, for instance, expose onchain history without every team rebuilding it from scratch.

DeFi and dApps

A lending protocol keeps the important state — who deposited what, who owes what — onchain, because it has to be trustless. But it reads prices from an oracle, and those prices originate offchain. The contract is onchain; the market data feeding it starts life somewhere else. If the price is wrong, the contract still executes, which is why the quality of that offchain feed matters so much.

NFTs

This one surprises people. When you buy an NFT, the ownership record is onchain, but the picture almost never is — it’s usually a file on IPFS or Arweave that the token points to through a tokenURI. That split is exactly why some early NFTs “lost” their images: the offchain file went offline while the onchain token kept working perfectly. Projects that care about longevity now lean on content-addressed or pay-once storage to avoid that fate.

DAOs and governance

Decentralized organizations show the split clearly. The binding vote — who voted, how much weight they carried, what passed — is recorded onchain so it can’t be quietly rewritten. The debate, proposals, and forum threads that lead up to it live offchain, because putting every comment on a blockchain would be slow and pointless to pay for.

Analytics and AI tools

Dashboards, tax tools, and AI agents read enormous amounts of onchain data, but they don’t query the chain live for every request. They index it into offchain databases first, then serve fast queries from there. If you’re building in that space, the best crypto data APIs handle most of that indexing for you.

Document and identity proofs

Some of the cleanest uses barely touch the chain at all. A platform can hash a contract, diploma, or audit report, store only that hash onchain, and keep the actual file offchain. Anyone can later re-hash the file and check it against the chain to prove it wasn’t altered — full verifiability, none of the storage cost.

How does offchain data get onto the blockchain?

Data moves across the onchain/offchain boundary in both directions, and a few systems do most of that work: oracles, indexers, and data-availability layers.

Oracles: the bridge for external data

An oracle is a service that fetches offchain data — a price, an exchange rate, a real-world outcome — and delivers it to a smart contract in a form the contract can trust (Chainlink). Blockchains can’t reach out to the internet on their own, so without oracles a contract has no idea what ether costs or whether a flight landed. That gap has a name: the oracle problem, the challenge of getting trustworthy outside data onto a chain without reintroducing a single point of failure.

The usual answer is a decentralized oracle network, where many independent sources report the same value and the contract acts on the aggregate. Oracles cover more than prices, too. They supply verifiable randomness for games and mints, proof-of-reserve data for stablecoins, and messages that move between chains.

The scale here is large. As of May 2026, Chainlink alone reported around $110 billion in total value secured across its data feeds and cross-chain services (crypto.news). That figure exists because so much onchain value depends on offchain data arriving correctly.

Indexers: reading onchain data at scale

Traffic flows the other way too. Raw onchain data is awkward to query — it’s arranged for verification, not for questions like “show me this wallet’s entire trading history.” Indexers solve that by reading the chain, reshaping the data, and storing the result in fast offchain databases that apps can query in milliseconds. Widely used indexing protocols and hosted data APIs both do this, and it’s why most dashboards and wallets feel instant even though the underlying chain is not.

Rollups and blob data

Layer-2 rollups run most of their activity offchain, then post compressed data back to Ethereum so anyone can reconstruct it. Since 2024, they use “blobs” — chunks of about 128 KB that Ethereum stores only temporarily, deleting them after roughly 18 days (ethereum.org). It’s a deliberate middle ground: available long enough to keep the rollup honest, cheap enough that fees stay low. For the mechanics, our guides on EIP-4844 proto-danksharding and modular blockchains go deeper.

Blobs make the point nicely. Onchain versus offchain isn’t always a clean binary — some data is fully onchain, some fully off, and some sits in between for a fixed window.

What are the main trade-offs and risks?

Each side fails in its own way, and the failures are worth knowing before you design around them.

Offchain data’s weakness is availability and trust. If the server goes down, the IPFS pin expires, or the oracle feeds a bad number, the onchain part is left pointing at something broken or wrong. A contract will act on a corrupted price without hesitation, because it can’t tell the difference — and attackers know it. A well-known class of DeFi exploit works by briefly distorting the price an oracle reports, then draining a protocol that trusted it.

Onchain data’s weakness is the flip side of its strength. It’s permanent and public, so a mistake is permanent and public too. Post private information or a buggy contract, and you can’t quietly take it back — you can only work around it. And you pay for every byte, up front and forever.

Privacy deserves its own line. Because onchain data is public by default, anything you write can be read, indexed, and linked back to an address by anyone. Sensitive data almost never belongs on a public chain in plaintext, which is another reason the hash-and-pointer pattern exists.

Practically speaking, the safest designs assume both can fail. Keep the trust-critical record onchain, store a hash so you’d notice if the offchain copy changed, and use more than one source for anything an oracle feeds you.

Choosing what goes where

The onchain vs offchain question isn’t really technical trivia — it’s a design decision you make on every project. A quick way to sort it: if the data must be trustless, verifiable, and permanent, it goes onchain — ownership, balances, settlement, the rules of a contract. If it’s large, private, or changes constantly, it stays offchain — media, documents, analytics, raw feeds. When you need both low cost and integrity, keep the content offchain and anchor a hash onchain.

Get that split right and you get the best of both — the guarantees of a blockchain without paying to store a video on it. Whether you’re reading onchain history or pulling offchain market data, the real work is usually about accessing both cleanly, not choosing one over the other.

FAQ

Do you pay gas to read onchain data?

No. You only pay gas to write data or change state. Reading is free at the protocol level, and most apps read through a provider or explorer rather than querying the chain themselves.

Can anything ever be removed from a blockchain?

Practically, no. Confirmed onchain data is replicated across the network and treated as immutable. You can add a new transaction that corrects or overrides an old one, but you can’t erase the original.

Is offchain data less secure than onchain data?

Not inherently — it has a different trust model. Offchain data’s safety depends on who hosts it and whether a hash anchors it, while onchain data secures itself through the whole network. Weak offchain handling is a design choice, not a rule.

What’s the difference between onchain and offchain transactions?

An onchain transaction is settled and recorded on the blockchain itself, so the whole network validates it. An offchain transaction happens outside the base chain — on a Layer 2, a payment channel, or between parties — and may touch the chain only later or in summarized form. Onchain means slower and costlier but fully settled; offchain means faster and cheaper but with extra trust or a later settlement step.

How can I tell whether a project keeps its data onchain or offchain?

Look at where value and ownership sit versus where media and feeds sit. Balances, token ownership, and contract logic are almost always onchain; images, documents, and external prices are usually offchain and referenced by a hash or URL.

Does storing data offchain make an app centralized?

Not necessarily. A centralized server does add a single point of control, but decentralized options like IPFS and Arweave spread data across many independent hosts. What matters is whether the data is content-addressed and anchored by an onchain hash, so users can verify it hasn’t changed regardless of who serves it.

Onchain or offchain storage — which should I use?

Usually both. Use onchain storage for the few things that must be trustless and permanent, and offchain storage for everything large, private, or fast-changing, joined back to the chain with a hash.