ar.io Logoar.io Documentation
ARIO Deploy

Incremental uploads

--incremental makes a redeploy pay only for the files that actually changed.

Arweave storage is permanent, so re-uploading byte-identical files buys nothing. Build tools content-hash their output, so between two deploys of a real site only a couple of entry chunks change — everything else is already on chain and can be referenced by its existing transaction id in the path manifest.

ario-deploy deploy --wallet ./wallet.json --incremental

Measured on a 1,229-file static docs site, redeployed from a fresh CI runner with no local cache and --compress gzip: 1,228 files were found on chain and one was uploaded (2.5 MiB), where a cold deploy uploaded 33 MiB.

How it works:

  1. Every file in the folder is hashed (SHA-256).
  2. Each file is looked up in the local dedupe cache, and then — for anything the cache cannot answer — among your own past uploads on chain.
  3. Only the remainder is uploaded, and each transaction id reaches the cache as it lands — on the leading edge, then coalesced onto a 500 ms trailing timer, and flushed on SIGINT/SIGTERM so Ctrl-C does not lose files you have already paid for. SIGHUP and SIGBREAK are not handled, so a closed terminal or a dropped SSH session can still lose the current batch; CI is covered, since GitHub Actions cancels with SIGINT then SIGTERM.
  4. The manifest is assembled from the remembered ids plus the new ones.

Why the on-chain lookup matters: every uploaded file carries a File-SHA256 tag, which makes it findable again from nothing but the bytes on disk. That is what a CI job needs. CI runs from a fresh checkout, so .ario-deploy/transaction-cache.json is often missing or stale even with actions/cache restoring it — and without the on-chain lookup every redeploy pays for the whole bundle again.

The tag invariant: a data item's id covers its tags, so a tag whose value changes between deploys — a commit SHA above all — moves every file's id on every deploy and defeats deduplication. The failure is silent: the upload succeeds, the manifest is correct, and the bill doubles. In incremental mode files therefore carry only deploy-invariant tags (App-Name, Content-Type, File-SHA256, plus Content-Encoding when compressed), and the GIT-HASH provenance tag rides on the manifest instead, which is rewritten every deploy anyway. The tag set is asserted in code, so a future addition fails loudly rather than quietly costing money.

Reuse is keyed on content type as well as content. Two files with identical bytes served under different types — a.json and b.txt — stay two uploads, because a gateway serves whatever Content-Type the data item carries and collapsing them would serve one of them as the other. Cache entries written in incremental mode are therefore keyed \<sha256\>|\<mime-type\> (plus |\<encoding\> when compressed); entries written by a plain (non-incremental) run stay keyed on the bare hash, so switching a project to --incremental re-uploads once and is cheap from then on.

What it trusts: only your own wallet's past transactions, matched on the 43-character address a gateway indexes an owner as — derived locally as base64url(sha256(publicKey)), which is correct for all five signer types. Every result is then re-checked here against the owner and the content type the gateway itself reports, because the owners filter is applied by whichever host --incremental-gateway names, and a wrong id would land in both the permanent manifest and the local cache.

Limits and caveats:

  • Lookups are batched. Hashes are sent 100 per GraphQL request, because gateways cap the size of a query (an ar.io gateway refuses ~1,100 hashes with "Max query size exceeded"). A site of any size is covered; each batch is paged until its files are accounted for, up to 20 pages.
  • The credits pre-flight prices only what will be sent: the files still to upload plus an estimate of the manifest, which is uploaded on every deploy. A fully reused redeploy is priced at the manifest alone.
  • Gateway GraphQL indexing lags an upload by a few minutes. Two machines deploying the same new file at the same moment can each pay for it. It costs a fraction of a cent and never produces a wrong manifest.
  • A gateway that is slow, unreachable or erroring costs reuse, not correctness. Requests that fail transiently (HTTP 429 or 5xx, a timeout, a network error) are retried twice with a short backoff. A batch that still fails costs only its own files, which are uploaded again, and the run says how many batches it could not look up. If no batch can be looked up at all, the run warns and uploads everything the local cache does not already hold.
  • A doomed deploy takes longer to say so. Every queued upload settles before a failure is reported, so a systemic failure (bad credentials, exhausted credits) on a very large folder surfaces at the end rather than immediately. The same uploads were always attempted, so the bill is unchanged; the alternative stranded ids that had been paid for and never written down.
  • Ignored for --deploy-file. Reuse works through the manifest, and a single file has no manifest. The run warns rather than silently doing nothing.
  • Cache entries are keyed differently in each mode, so a project that toggles --incremental on and off stores up to two entries per file against the shared --dedupe-cache-max-entries cap: \<sha256\> (or gzip:\<sha256\> when compressed) without it, and \<sha256\>|\<mime-type\> (or \<sha256\>|\<mime-type\>|gzip) with it.

Notes:

  • Off by default. Nothing changes for an existing pipeline until you pass the flag.
  • Refused alongside --no-dedupe or --dedupe-cache-max-entries 0, which ask for the opposite.
  • Works with --compress: each file's File-SHA256 is the hash of the file on disk, and a compressed upload also carries Content-Encoding, so a lookup only ever reuses an upload made with the same encoding. Turning compression on or off uploads each file once more, then reuse resumes.
  • The lookup uses https://turbo-gateway.com/graphql by default, where uploads made through Turbo are indexed within minutes (about 5-7 in our measurements), before they are bundled into a block. Override it with --incremental-gateway — for example when uploading through another bundler with --uploader.

How is this guide?