What Is an RPC Endpoint? How Requests, Rate Limits, and Latency Work

An RPC endpoint is a URL your application sends requests to in order to read data from a blockchain or broadcast a transaction to it. RPC stands for Remote Procedure Call: your code calls a function, that function runs on a machine somewhere else, and you get the result back. On a blockchain, that machine holds a copy of the ledger, and the endpoint is the address where it listens.

This guide covers the four things developers ask about most. What an endpoint actually is, how RPC requests are structured, why providers put a rate limit on them, and what makes one endpoint slower than another.

What Is an RPC Endpoint?

An RPC endpoint is a URL, usually HTTPS or WebSocket, where an application sends JSON-formatted requests and receives blockchain data in return. A typical one looks like https://eth.example.com/YOUR_API_KEY, where the key identifies your account and tracks your usage.

The endpoint is not the same thing as the server behind it. That server stores blockchain state, validates data, and executes your request. The endpoint is just the door you knock on. One provider can route many endpoints to a shared pool of machines, and a single application can point at several endpoints across different networks at once.

Because JSON-RPC is a shared standard, endpoints are largely interchangeable between providers. Auston Bunsen, co-founder of QuickNode, put it plainly in an interview with Sacra: “I can go from Alchemy to Infura to QuickNode relatively quickly, unless I’m using one of their sort of custom APIs.” Swapping providers is usually a change of URL and key, not a rewrite. The exception is when you have built on a provider’s proprietary methods, which do not carry over.

HTTP or WebSocket?

Most requests go over HTTP(S). You send one request, you get one response, the connection closes. That works well for occasional reads and for sending transactions.

WebSocket keeps a single connection open so the server can push updates to you, such as new blocks, pending transactions, or contract events, without you asking again and again. For anything event-driven, a persistent connection avoids the waste of polling the same question on a loop. The right transport depends on whether you are pulling data on demand or reacting to things as they happen.

How Do RPC Requests Work?

RPC requests to a blockchain are small JSON objects, almost always built on the JSON-RPC 2.0 standard. Ethereum and every EVM-compatible network use it, which is why the same request shape works across so many chains (Ethereum JSON-RPC docs).

Four fields do the work:

  • jsonrpc — the version string, always "2.0"
  • method — the function you want to run, such as eth_getBalance
  • params — the arguments that method needs
  • id — a number or string you pick so you can match the reply to the request

A request to read an account balance looks like this:

json

{
  "jsonrpc": "2.0",
  "method": "eth_getBalance",
  "params": ["0x742d35Cc6634C0532925a3b844Bc454e4438f44e", "latest"],
  "id": 1
}

The server answers with a matching object that contains either a result or an error, never both (JSON-RPC 2.0 spec). When something goes wrong, the error carries a numeric code: -32601 means the method does not exist, -32602 means the parameters were wrong, and -32700 means the JSON itself could not be parsed. Reading those codes turns a vague “it failed” into a specific fix.

Methods fall into two groups. Read methods ask for existing data and cost nothing in gas. Write methods submit something to the network.

MethodWhat it doesType
eth_blockNumberReturns the latest block numberRead
eth_getBalanceReturns an address balanceRead
eth_callRuns a contract function without sending a transactionRead
eth_getLogsFetches event logs matching a filterRead
eth_gasPriceReturns the current gas priceRead
eth_sendRawTransactionBroadcasts a signed transactionWrite

Those are Ethereum and EVM method names. Other networks expose their own, so a Bitcoin or Solana endpoint will not answer eth_getBalance.

Batching requests

If you need several pieces of data at once, you can send an array of request objects in a single call and get an array of responses back. Batching cuts the number of round trips, which matters more than it first appears once latency enters the picture, covered further down.

What Is RPC Rate Limiting?

An RPC rate limit is the cap a provider puts on how much you can ask for in a given window. Cross it and the endpoint starts refusing requests, usually with the HTTP status 429 Too Many Requests.

The limit exists because a server has finite resources and is shared among customers. One application firing an unbounded loop of heavy queries would slow the service down for everyone else on that infrastructure. Rate limiting keeps one noisy tenant from starving the rest.

Providers measure usage in a few different ways, and the model matters as much as the headline number:

ModelWhat it countsWhere you see it
Requests per second (RPS)How many calls you send each secondCommon on shared plans
Compute unitsA weighted cost per method, since some calls are heavier than othersAlchemy and others
Monthly request quotaTotal calls allowed per billing cycleMost providers

The compute-unit model reflects something real about RPC: not all calls are equal. Alchemy’s documentation notes that “some queries are lightweight and fast to run (e.g., eth_blockNumber) and others can be more intense (e.g., large eth_getLogs queries),” and it prices them accordingly, with one of its methods listed at 100 compute units and 10 compute units per second (Alchemy docs). A plan advertised as a flat number of requests can run out far sooner if your calls are the expensive kind.

NOWNodes uses monthly request quotas on its shared plans, which currently start at 100,000 requests and rise into the tens of millions on higher tiers. Its free public endpoints carry a low fixed rate limit and are meant for prototyping and testing rather than production traffic. Dedicated nodes sit at the other end: NOWNodes states its dedicated nodes have no predefined RPS limit, with throughput bounded only by the hardware allocated to that node.

How to handle rate limits

A few habits keep an application stable when it bumps against a limit:

  • Retry with exponential backoff instead of hammering the endpoint again right away.
  • Cache data that rarely changes, such as a token’s decimals or a finalized block.
  • Batch related reads so one request does the work of several.
  • Move to a higher tier or a dedicated endpoint once steady traffic outgrows a shared quota.

The goal is to stay under the ceiling by design, not to react every time a 429 comes back.

What Is RPC Latency?

RPC latency is the time between sending a request and getting the response back, the full round trip, measured in milliseconds. It is not the same as block time or confirmation time. A fast endpoint can still hand you data that is a few seconds old if the chain itself produces blocks slowly, so low latency means a quick answer, not necessarily the freshest state.

Latency matters wherever timing is competitive. A trading bot acting on a price several hundred milliseconds late is working from stale information, and a wallet that takes a full second to show a balance feels broken. For a backend crunching analytics overnight, the same delay is irrelevant.

Several things add up to the number you see:

  • Distance. A request crossing an ocean carries round-trip delay that a nearby server simply does not.
  • Server load. An overloaded or under-provisioned machine answers more slowly.
  • Request complexity. eth_blockNumber returns almost instantly, while a wide eth_getLogs scan across thousands of blocks is real work.
  • The connection itself. DNS lookup, TCP setup, and the TLS handshake all add milliseconds before your request even arrives.

You reduce latency by shortening both the path and the work. Pick a provider with servers near your users, reuse connections instead of opening a fresh one per call, keep heavy queries narrow, and subscribe to events over WebSocket or gRPC streaming rather than polling in a tight loop. NOWNodes publishes response times around 0.2 seconds and latency under 200 milliseconds for key US and European regions, served from geo-balanced infrastructure; those are its stated figures for those regions, not a worldwide guarantee.

Latency vs throughput

The two get mixed up often. Latency is how fast one request completes. Throughput is how many requests you can complete per second. An endpoint can have low latency and still throttle you on throughput, or handle high volume while each individual call is slow. You care about latency for responsiveness and throughput for scale, and a shared plan’s rate limit is really a throughput ceiling.

Connecting to an RPC Endpoint

Getting an endpoint of your own takes a few steps and no server maintenance on your side:

  1. Create an account with a provider and generate an API key.
  2. Copy the endpoint URL for the network you need.
  3. Add the URL and key to your application, wallet, or library configuration.
  4. Send a basic request such as eth_blockNumber to confirm the connection.

A multi-network provider such as NOWNodes covers this across 120+ chains from a single account, so an application that needs Ethereum, BNB Smart Chain, Bitcoin, and a few others does not have to stitch together one endpoint per chain from separate vendors (NOWNodes network directory). The trade-off is the familiar one: shared endpoints are cheaper and quick to start but carry quotas, while dedicated endpoints cost more and drop the predefined limits.

Conclusion

An RPC endpoint is the connection point between your application and a blockchain. You send JSON-RPC requests to a URL, a server runs them, and you get data or a transaction receipt back. Requests follow a simple four-field format, rate limits protect shared infrastructure and come in request-count or compute-weighted forms, and latency is the round-trip time that decides whether your app feels fast or sluggish.

For most projects the real question is not how to run the infrastructure but which endpoint to point at. Match the plan to your actual traffic, watch how heavy your calls are rather than just how many, and choose a region close to the people using your app.

FAQ

What Does RPC Stand For?

Remote Procedure Call. Your application calls a function that runs on a remote server and returns the result, which on a blockchain means reading ledger data or broadcasting a transaction.

Is an RPC Endpoint the Same as an API?

An RPC endpoint is one kind of API. It exposes blockchain functions through the JSON-RPC standard at a single URL, rather than spreading them across the many resource paths a REST API would use.

Are Free Public RPC Endpoints Safe for Production?

They are fine for testing and small projects, but their shared rate limits and lack of guarantees make them a poor fit for production traffic. A keyed plan or a dedicated endpoint gives you far more predictable behavior.

Why Am I Getting a 429 Error?

A 429 means you went over the endpoint’s rate limit. Slow your request rate, add retry-with-backoff, cache repeated reads, or move to a plan with a higher ceiling.

Can One API Key Work Across Several Blockchains?

With a multi-network provider, yes. You generate a key once and switch the endpoint URL per network, since the JSON-RPC format is shared across EVM chains. Non-EVM chains expose their own method sets.