Logo
New RPC users get 35% off their first monthView the offer
OnFinality Learn
Infrastructure & Operations14 min read

Polygon Flat Files: Bulk Historical Chain Data Without RPC Limits

Use Polygon flat files to backfill historical block, transaction, and log data in bulk instead of exhausting JSON-RPC request budgets.

TL;DR

Polygon flat files are periodically produced, immutable exports of raw block, transaction, and log data stored in object storage and addressed by block range, not a live query endpoint. They exist because backfilling millions of blocks over JSON-RPC would require millions of eth_getBlockByNumber or eth_getLogs calls, consuming enormous request-unit budgets and triggering HTTP 429 responses. Flat files transfer the same data as a handful of large sequential object downloads. Use live JSON-RPC for current state and recent blocks, an archive node for historical state queries like eth_call at a past block, and flat files for historical block, transaction, and log data you will process in bulk offline. Flat files are only as fresh as the last export, are network-specific, and their exact bucket layout and format vary by provider and must be read from that provider's documentation.

What Polygon Flat Files Actually Are

Polygon flat files are periodically produced, immutable files containing raw block, transaction, and log data for the Polygon chain. They are stored in object storage, commonly an S3-compatible bucket, and are addressed by block range rather than queried interactively. Each export is a snapshot: once written, the file does not change, and a new file appears for the next range. This is fundamentally different from a JSON-RPC endpoint, which responds to live requests for current state and recent blocks.

The format is typically CSV or Parquet-style columnar exports, though the exact schema and compression vary by provider. Because they are files, you download them with standard object-storage tooling and parse them locally or in your own pipeline. They are not a replacement for RPC; they are a bulk data-transfer surface for historical analysis, indexing, and backfills. For a broader overview of data-access surfaces, see Accessing historical blockchain data.

  • Immutable, periodically produced exports of block, transaction, and log data.
  • Stored in object storage (e.g., S3-compatible bucket) and addressed by block range.
  • Not a live query endpoint; no eth_call or eth_getLogs semantics.
  • Format and layout vary by provider; always read the provider's documentation.

Why Flat Files Exist: The RPC Backfill Problem

An indexer that needs to backfill millions of blocks would otherwise issue millions of eth_getBlockByNumber or eth_getLogs calls. Each call consumes request units, and sustained high-volume reads quickly exhaust rate limits, producing HTTP 429 responses. The JSON-RPC 2.0 specification defines a request-response protocol optimized for interactive queries, not bulk sequential extraction. The Ethereum JSON-RPC specification documents the block, transaction, and log methods that a bulk extract replaces.

Flat files invert the model: instead of millions of small requests, you perform a handful of large sequential object downloads. The data volume is the same, but the request count drops by orders of magnitude, and you avoid the per-request overhead and rate-limit pressure. This is why flat files are the preferred surface for historical backfills, while RPC remains the right tool for current state and recent blocks. For a deeper look at rate-limit mechanics, see Polygon RPC 429 and rate limits.

  • Millions of RPC calls consume request units and trigger 429s.
  • Flat files transfer the same data as a few large object downloads.
  • RPC is optimized for interactive queries, not bulk sequential extraction.
  • Use flat files for backfills; use RPC for current state and recent blocks.

Three Polygon Data Paths: Live RPC, Archive Node, and Flat Files

Polygon exposes three distinct data-access surfaces, and choosing the wrong one is the most common source of wasted budget and failed backfills. Live JSON-RPC serves current state and recent blocks; it is the right surface for wallets, dApps, and real-time monitoring. An archive node serves historical state queries that need eth_call at a past block, such as checking a token balance at a specific height. Flat files serve historical block, transaction, and log data you will process in bulk offline.

The distinction matters because archive nodes and flat files are not interchangeable. An archive node can answer state queries at any block, but querying millions of blocks through it still incurs per-request costs and rate limits. Flat files cannot answer state queries at all, but they deliver historical block and log data far more efficiently. For a detailed comparison of node types, see Archive node vs full node and Polygon archive nodes and historical RPC.

  • Live JSON-RPC: current state, recent blocks, real-time dApps.
  • Archive node: historical state queries (eth_call at a past block).
  • Flat files: historical block, transaction, and log data for bulk offline processing.

Worked Decision Table: Choosing the Right Surface

Use this table to map a task to the best data surface. The goal is to avoid using a live RPC endpoint for bulk historical extraction, and to avoid using flat files for anything that requires state or the chain tip. The table reflects documented behavior; provider-specific details such as bucket layout and export cadence vary and should be confirmed in the provider's documentation.

When a task spans multiple surfaces, split it: use flat files for the historical bulk, and RPC for the recent tail and any state lookups. This hybrid pattern keeps request-unit consumption low while preserving correctness for state-dependent logic.

  • Backfill millions of blocks of transactions/logs → Flat files → Avoids millions of RPC calls and 429s.
  • Query current token balance → Live JSON-RPC → State is only available via RPC or a database.
  • Check historical balance at block N → Archive node → Flat files do not contain state.
  • Stream recent blocks for a dashboard → Live JSON-RPC → Flat files lag behind the chain tip.
  • Build a historical analytics dataset → Flat files → Bulk sequential downloads are efficient.
  • Resolve a reorg near the tip → Live JSON-RPC → Flat files are immutable snapshots, not live.

How Flat Files Are Laid Out and Enumerated

Flat files are organized by block range and stored under a prefix in an object-storage bucket. To enumerate them, you list the bucket for your target block range, download the relevant keys, decompress if needed, and stream-parse the contents rather than loading everything into memory. The exact key structure, file naming, and compression vary by provider, so always read the provider's documentation for the authoritative layout. Polygon's developer documentation at docs.polygon.technology describes data-access surfaces and RPC endpoints, while flat-file bucket specifics are provider-defined.

A typical workflow is: determine the block range you need, list the bucket prefix that covers that range, download each object, and parse it line by line or row by row. Because files are immutable, you can cache them locally and re-process without re-downloading. For large ranges, process files in parallel up to your bandwidth and memory limits, but keep the parse streaming to avoid out-of-memory errors.

  • List the bucket prefix for your block range to discover available files.
  • Download by key; files are immutable and can be cached locally.
  • Decompress if the provider stores compressed exports.
  • Stream-parse rather than loading entire files into memory.
  • Confirm exact layout and format in the provider's documentation.

Runnable Example: Listing and Streaming a Flat-File Object

The following Node.js example lists an S3-compatible flat-file prefix for a block range, downloads one object, and streams-parses it to count transactions. It uses the AWS SDK v3, which works with any S3-compatible endpoint. Replace the bucket, prefix, and credentials with your provider's values. The code assumes newline-delimited JSON or CSV; adjust the parser to match the provider's format.

This example demonstrates the core pattern: enumerate, download, stream-parse. It does not load the entire file into memory, which is essential for large exports. For a full production pipeline, add error handling, retries, and parallel downloads within your resource limits.

const { S3Client, ListObjectsV2Command, GetObjectCommand } = require('@aws-sdk/client-s3');
const { createGunzip } = require('zlib');
const readline = require('readline');

const s3 = new S3Client({
  region: 'us-east-1',
  endpoint: 'https://your-s3-compatible-endpoint',
  credentials: { accessKeyId: process.env.AWS_ACCESS_KEY_ID, secretAccessKey: process.env.AWS_SECRET_ACCESS_KEY },
  forcePathStyle: true,
});

async function listFlatFiles(bucket, prefix) {
  const res = await s3.send(new ListObjectsV2Command({ Bucket: bucket, Prefix: prefix }));
  return (res.Contents || []).map((o) => o.Key);
}

async function streamCountTransactions(bucket, key) {
  const res = await s3.send(new GetObjectCommand({ Bucket: bucket, Key: key }));
  const stream = res.Body.pipe(createGunzip());
  const rl = readline.createInterface({ input: stream, crlfDelay: Infinity });
  let txCount = 0;
  for await (const line of rl) {
    if (!line.trim()) continue;
    const row = JSON.parse(line);
    if (row.transactionHash) txCount++;
  }
  return txCount;
}

(async () => {
  const bucket = 'your-flat-file-bucket';
  const prefix = 'polygon/blocks/19000000/'; // adjust to provider layout
  const keys = await listFlatFiles(bucket, prefix);
  console.log('Found files:', keys.length);
  if (keys.length) {
    const count = await streamCountTransactions(bucket, keys[0]);
    console.log('Transactions in', keys[0], ':', count);
  }
})();

Contrast: The Equivalent RPC Loop and Its Cost

To appreciate the difference, consider the RPC equivalent of the previous example. Backfilling the same block range over JSON-RPC would require one eth_getBlockByNumber call per block, plus additional calls for logs if needed. For a range of one million blocks, that is one million requests, each consuming request units and counting against rate limits. The Ethereum JSON-RPC specification documents these methods, and the JSON-RPC 2.0 specification defines the request-response model that makes per-block reads inefficient for bulk extraction.

The following snippet shows the RPC loop pattern. It is correct for small ranges or recent blocks, but it does not scale to millions of blocks without hitting HTTP 429 responses. Use it for the recent tail, not for historical backfills.

const { ethers } = require('ethers');

async function rpcBackfill(providerUrl, fromBlock, toBlock) {
  const provider = new ethers.JsonRpcProvider(providerUrl);
  let txCount = 0;
  for (let n = fromBlock; n <= toBlock; n++) {
    const block = await provider.send('eth_getBlockByNumber', ['0x' + n.toString(16), false]);
    if (block && block.transactions) txCount += block.transactions.length;
  }
  return txCount;
}

// For a large range, this loop will consume enormous request units
// and likely trigger HTTP 429 responses. Use flat files instead.

Reproducible Results Table: Measure Your Own Pipeline

To make an informed decision, measure your own pipeline against your own endpoint and provider. The table below is a template: fill it with your own measurements. Do not rely on vendor benchmarks or third-party numbers; the variables that matter (block range, file size, network bandwidth, provider rate limits) are specific to your environment. Run the same logical task over flat files and over RPC, and record the results.

Use the table to compare wall time, bytes transferred, and HTTP 429 counts. The goal is not to prove one surface is always faster, but to quantify the tradeoff for your workload. For pricing implications of RPC request units, see RPC pricing.

  • Task: describe the operation (e.g., count transactions in blocks 19,000,000–19,100,000).
  • Method: flat files or RPC loop.
  • Blocks covered: the exact range.
  • Wall time: total elapsed time in seconds.
  • Bytes: total data transferred.
  • HTTP 429s: number of rate-limit responses encountered.

Limitations and Tradeoffs of Flat Files

Flat files are only as fresh as the last export. They cannot serve the chain tip, and they cannot answer state queries such as eth_call at a past block. They are network-specific: a Polygon flat file applies to Polygon, not to another chain. The schema is stable only within a documented version; if the provider changes the format, your parser must adapt. The exact bucket layout, file naming, and compression vary by provider and must be read from that provider's documentation.

You still need RPC or a database for state, mempool, and anything younger than the newest export. Flat files are a bulk historical data surface, not a general-purpose query engine. For state-dependent logic, pair flat files with an archive node or a live RPC endpoint. For a comparison of node types, see Archive node vs full node.

  • Only as fresh as the last export; cannot serve the chain tip.
  • Network-specific; not portable across chains.
  • Schema-stable only within a documented version.
  • Bucket layout and format vary by provider.
  • No state, mempool, or recent-block coverage.

Troubleshooting Common Flat-File and RPC Issues

When a flat-file pipeline fails, the cause is often a mismatch between the provider's layout and your assumptions. If listing returns no keys, verify the prefix and block range against the provider's documentation. If parsing fails, check the compression and format; a gzip stream fed to a CSV parser will produce errors. If downloads are slow, check your bandwidth and consider parallel downloads within your limits. If you see HTTP 429 responses, you are likely still using RPC for bulk extraction; switch to flat files for the historical range.

For RPC-specific issues, such as rate limits and 429 handling, see Polygon RPC 429 and rate limits. For archive-node state queries, see Polygon archive nodes and historical RPC. For a general RPC guide, see the Polygon RPC guide (RPC Assistant).

  • No keys listed: verify prefix and block range in provider docs.
  • Parse errors: confirm compression and format before parsing.
  • Slow downloads: check bandwidth; parallelize within limits.
  • HTTP 429s: you are likely using RPC for bulk extraction; switch to flat files.
  • State queries failing: flat files do not contain state; use an archive node.

Next Steps: Integrating Flat Files Into Your Stack

Start by identifying the historical block range you need and the provider whose flat files cover it. Read that provider's documentation for bucket layout, format, and export cadence. Build a small pipeline that lists, downloads, and stream-parses one file, then scale to your full range. Measure your pipeline with the results table above, and compare against an RPC loop for the same range. Use flat files for the bulk historical load, and reserve RPC for the recent tail and state queries.

For network-specific details, see the Polygon network page. For API and service options, see the API service. For more guides, visit the OnFinality Learn hub. For pricing implications, see RPC pricing.

  • Identify the historical block range and the provider whose flat files cover it.
  • Read the provider's documentation for layout, format, and cadence.
  • Build a minimal list-download-parse pipeline, then scale.
  • Measure with the results table and compare against RPC.
  • Use flat files for bulk history; use RPC for the recent tail and state.

Never Worry about Infrastructure Again

OnFinality takes away the heavy lifting of DevOps so you can build smarter and faster.

Get Started