New Zealand Web Archive
New Zealand Web Archive
- providers.nlnz()
- provider: "nlnz"
- Index
- CDX
- Bodies
- listing only
- provider=all
- included
Operations
- Listsnapshots()
- Readunsupported
- Diffunsupported
Access
- Load
createArchive(providers.nlnz()) - Needs
nothing, no key or collection - Replay
Replay hides behind a bot check. Listing works, reading doesn't.
Load it
import { createArchive, providers } from "@agntn/archives";
const archive = createArchive(providers.nlnz());
const response = await archive.snapshots("natlib.govt.nz");
What it lists
snapshots asks ndhadeliver.natlib.govt.nz/webarchive/cdx and gets NDJSON back. It's pywb again, so a bare domain is a prefix. natlib.govt.nz lists the whole site, back to a front page from July 2004. A thousand rows come back in about three seconds. Not bad for the other side of the planet.
Rows carry status, mime and digest in _meta. length stays out on purpose. The index writes "0" there for almost every record, and a zero that means "no idea" is worse than nothing.
limit, from and to behave. limit=0 still means "everything" to pywb, so the provider never sends it.
Why can't it read a capture?
Because the replay sits behind Imperva. A script asking for /webarchive/<timestamp>id_/<url> gets a JavaScript challenge page. With status 200, of course. Headless Chromium does no better. Read that naively and you'd get a very confident capture of a bot check.
So content() doesn't try. It answers unsupported and says why, and providers.all() moves on to an archive that does serve bodies. The snapshot link points at the archive's own playback page, behind the same check. Whether it lets you in is between you and Imperva.
The CDX index passes the same CDN today. If Imperva ever starts challenging it too, listings fail as a malformed CDX record, not as an empty success.
Gotchas
- The index canonicalizes the scheme. Ask for
https://www.nzherald.co.nz/and the rows come back ashttp://. - Coverage is New Zealand. Mostly
.nz, harvested by the National Library. - Need the bytes of an NZ page? Find the date here, then read the nearest Wayback or Common Crawl capture of the same URL.
Where it lives
src/providers/nlnz.ts.