Vefsafn
Vefsafn
- providers.vefsafn()
- provider: "vefsafn"
- Index
- CDX
- Bodies
- read raw
- provider=all
- included
Operations
- Listsnapshots()
- Readcontent()
- Diffdiff()
Access
- Load
createArchive(providers.vefsafn()) - Needs
nothing, no key or collection - Replay
the playback page opens framed on this site
Load it
import { createArchive, providers } from "@agntn/archives";
const archive = createArchive(providers.vefsafn());
const response = await archive.snapshots("https://www.ruv.is/");
What it lists
snapshots asks vefsafn.is/cdx about one exact URL. The answer is NDJSON, one capture per line. Rows carry status, mime and digest in _meta. The index ignores the scheme, so http:// and https:// get the same rows.
Why does it need a limit of its own?
Because the index ignores yours. Ask for three rows and it sends every capture of the page. A news front page has tens of thousands. So the provider reads the stream line by line and hangs up at the limit. You get three rows. The archive stops sending megabytes nobody reads.
Reading bodies
Captures replay under id_ at the root: vefsafn.is/<timestamp>id_/<url>. The body is the original response. Recent crawls of busy sites are full of 301 rows. A read without an exact timestamp skips them and takes the newest capture that answered with a 2xx. Want the redirect itself? Pass its exact timestamp.
Gotchas
- Exact URL only.
ruv.ismeanshttp://ruv.is/. A prefix query on a big host sent nothing back in 25 seconds, so the provider never asks one. - Old ARC records spell the port out:
http://www.ruv.is:80/. Same page, different string. - It goes back further than you'd guess. The oldest row for
ruv.isis from December 1996. - Coverage is Icelandic. The National and University Library crawls
.isand sites about Iceland.
Where it lives
src/providers/vefsafn.ts.