finca

About finca

Specialty coffee has no shared list of who grew what. The same coffee appears at five roasters under five names; producer names are spelled inconsistently and often left out; lots vanish in eight weeks. finca reads roasters' own shops, pulls out who grew each coffee (producer, farm, washing station, importer, varietal, process, crop year) and links those names into one public graph, with the price per 100 g beside each coffee.

What the crawler does

The crawler reads roasters' own public product feeds and shop pages, nothing else. It honors robots.txt, waits between requests to the same host, and identifies itself with a descriptive User-Agent that names the project and the delist policy. Every crawl attempt is recorded, failures included, and each roaster's crawl health is public on the roaster directory.

What we record

Each crawl writes one timestamped observation per listing: title, listed price, bag weight, availability. Observations are never edited or deleted. Price history, time-to-sellout and restock behavior depend on that record staying exactly as it was seen.

What we publish

The facts read from each listing (producer, farm, washing station, importer, varietal, process, crop year) plus a price per 100 g. A price in another currency is converted at the exchange rate on the day it was seen, never today's rate, so a price series tracks coffee rather than currency. VAT-inclusive and sale prices are marked.

Product pages show the description the roaster wrote, labelled as theirs and linked back to their listing. Product photos load straight from the roaster's servers with a link back to the listing; they are not re-hosted.

What the confidence labels mean

Reading a listing and deciding which farm or producer a name refers to are separate steps, and every link carries a confidence score and a trail back to the listing it came from. Every link you see on a page has passed our confidence checks; we don't mark some links as more certain than others. Guesses that fail the bar aren't shown at all; they're held for review instead of being published.

The threshold is one central setting, so the whole site recalibrates together as real error rates become known. The claim that two roasters sell the same green lot is always an inference with its own confidence, never presented as certain.

What we don't do

Nothing here is for sale. No ranking, badge or slot on this site can be bought or influenced by a commercial relationship. Directories are alphabetical or chronological; nothing is featured. There are no affiliate links, ads, sponsorships or referral parameters; the only commercial link on a product page is the plain link to the roaster's listing.

The roaster's words stay the roaster's. The description on a product page is theirs and labelled as theirs; the facts we extract and link are ours.

Sold-out pages stay up. A sold-out coffee is part of the historical record, not a deleted page.

Free to use, not free to take. The graph is viewable by anyone; bulk reuse is a different conversation.

Errors, merges and permanence

Names get misspelled and records get merged. When two records turn out to be one producer, they merge, and the losing page address keeps working forever by forwarding to the winner. An address is never reused for a different producer or farm.

Reading the pages

Delisting

Delisting on request is immediate and no-argument. A one-line email to [email protected] (the contact in our robots.txt and the crawler's User-Agent) takes a roaster out of every public directory and marks their pages noindex.

Corrections and questions: [email protected].