Providers

Vefsafn

Iceland's national web archive. One exact URL per query, id_ replay for bodies, and a limit the provider enforces itself.
IDvefsafn04 / 11vefsafn.is
Provider

Vefsafn

  • providers.vefsafn()
  • provider: "vefsafn"
Index
CDX
Bodies
read raw
provider=all
included

Operations

Access

LoadcreateArchive(providers.vefsafn())
Needsnothing, no key or collection
Replaythe playback page opens framed on this site

Load it

ts
import { createArchive, providers } from "@agntn/archives";

const archive = createArchive(providers.vefsafn());
const response = await archive.snapshots("https://www.ruv.is/");

What it lists

snapshots asks vefsafn.is/cdx about one exact URL. The answer is NDJSON, one capture per line. Rows carry status, mime and digest in _meta. The index ignores the scheme, so http:// and https:// get the same rows.

Why does it need a limit of its own?

Because the index ignores yours. Ask for three rows and it sends every capture of the page. A news front page has tens of thousands. So the provider reads the stream line by line and hangs up at the limit. You get three rows. The archive stops sending megabytes nobody reads.

Reading bodies

Captures replay under id_ at the root: vefsafn.is/<timestamp>id_/<url>. The body is the original response. Recent crawls of busy sites are full of 301 rows. A read without an exact timestamp skips them and takes the newest capture that answered with a 2xx. Want the redirect itself? Pass its exact timestamp.

Gotchas

  • Exact URL only. ruv.is means http://ruv.is/. A prefix query on a big host sent nothing back in 25 seconds, so the provider never asks one.
  • Old ARC records spell the port out: http://www.ruv.is:80/. Same page, different string.
  • It goes back further than you'd guess. The oldest row for ruv.is is from December 1996.
  • Coverage is Icelandic. The National and University Library crawls .is and sites about Iceland.

Where it lives

src/providers/vefsafn.ts.