Skip to content
The Kennel Bench

One kennel's drafts, measured against the standard.

BOOK-004BOOK

What a capture keeps, and what it drops

Every sheet on this bench rests on a Wayback Machine capture. None of them explains what a capture actually preserves, and what it quietly loses. This one does.

Filed by The Kennel Bench · Entry checked on · 4 min read · Soul Guardian Bulldogs, Greece

Eleven printed bulldog portraits laid in a loose grid on a workbench

BOOK-004 · pinned 06/10/2026

Eleven printed photographs laid in a grid on a workbench.

Fourteen earlier sheets on this bench each cite a web.archive.org address and each, in its own way, notes what its particular capture does not keep. None of them explains the machine itself. A marquee file of seasons is itself a record of dated dispatches; this sheet reads the machine that dates and keeps every dispatch this bench has filed so far: the Wayback Machine, run by the Internet Archive, a nonprofit that has kept it open to the public since October 2001.

What a capture actually is

A Wayback Machine capture is a snapshot: a crawler visits a page's address at a given moment, downloads what is publicly reachable there, and stores it with a timestamp. Every source cited on this bench, dogs.html in 2012, pets.html, the standard page, the show page, is one such snapshot, not a continuously updated mirror of the original site. That is why every sheet on this bench reads one dated version of a page rather than its whole history: the capture is the only version the archive chose, or was able, to keep.

Why some pages have many captures and others have one

A site's pages get crawled on no fixed schedule of the owner's choosing. Some addresses were captured repeatedly across 2010 to 2013; the earlier sheet on this kennel's available-puppies page found four separate addresses for roughly the same content, which this bench read as a page the kennel kept revising. Others, like the pets page, survive as a single capture. The difference is not a judgment on which page mattered more to the kennel; it reflects how often the Internet Archive's own crawlers, or a third party's submitted crawl, happened to pass by that particular address.

What a capture keeps well

A capture keeps text and most ordinary images reliably, which is why this bench can quote filenames like boubou.jpg or lilly_.html with confidence. It keeps a page's structure, its directory address, and its publication date as that one crawl recorded it. For a site built by hand in the early 2010s, mostly static HTML pages and photographs, this is close to a full preservation: the kennel's site was exactly the kind of simple, file-based structure a crawler captures cleanly.

What a capture drops

Published accounts of the Wayback Machine's limits name several gaps that matter to how this bench reads its sources. A crawler cannot fully capture interactive features, scripts that build a page's content after it loads, or anything that requires a visitor to submit a form before content appears. A page with no other page linking to it, an "orphan" address, may never be crawled at all, however it might have reached visitors through, say, a guestbook, an email, or a printed card. None of this means a capture lies; it means a capture is a photograph of what a crawler could see from the outside, not a copy of everything the original page's owner once published.

Why this matters for a kennel's own record

Every "what the record does not keep" line on the other thirteen sheets, the show page's results, the pets page's captions, the standard's exact wording, is this same limitation working in a specific case. A result typed into a now-unreadable table cell, a caption rendered by a script the crawler never ran, a page nobody linked to and so nobody's crawler ever found: these are not failures of this bench's reading, they are the shape of what a capture was built to do and not do. Reading the machine once, here, lets every other sheet's silence be read for what it actually is, a known limit of the tool, rather than a mystery each time it recurs.

The machine's own scale

The Internet Archive reports having archived more than a trillion web pages, a scale that makes a single small Greek kennel's site one address among an enormous number, kept for the same reason most of the rest were: a crawler passed, the page was public, and nothing told the machine to skip it. That scale is also why a single hand-built kennel site from 2010 survives at all into 2026: not because anyone judged it worth preserving, but because the machine that keeps the record does not need a reason to keep an ordinary page, only an address and a moment it was reachable.

How the bench files it

The Wayback Machine is filed here as the bench's own tool, not as another dog or another page: a nonprofit's crawler, open to the public since 2001, that keeps what it can see from outside a page and drops what it cannot, the same honest limit every other sheet on this bench has already been working within.

Neighbouring entries

A dog collar laid on a folded wool blanket in the corner of a quiet room

BOOK-002BOOK

The in-memory page

The page that mourned the kennel's animals

· 3 min

A bulldog, a kitten and a parrot sharing a sunlit living-room floor

BOOK-001BOOK

The other animals

The pets page and its eight photographs

· 3 min