Methodology
How PhotoDataLab analyzes photography metadata
PhotoDataLab turns selected public-photo metadata into versioned observational datasets. Sample size, contributor diversity, field coverage, and uncertainty stay visible beside every result.
Observations, not universal advice
Results use language such as “observed” and “most frequently observed.” A common value is not automatically an optimal value, a recommendation, or a measure of market popularity.
Scope and exclusions
Dataset version 2 contains 166,036 photo observations from 1,063 photographers across 5,537 public albums. Photo files are not copied or embedded. Unnecessary personal information and GPS coordinates are excluded from the public dataset.
Camera and lens identities
Raw EXIF make, model, and lens strings are cleaned by deterministic versioned normalizers. Alias records preserve how a value appeared while mapping proven formatting and manufacturer-prefix variants to one canonical identity. Ambiguous generic strings are not presented as specific products.
Genre classification
Genres are inferred from explicit keywords in public album names by classifier
album-keywords-v1. This is a reproducible heuristic, not a human-verified label for every photo.
Contributor balance
Photo-weighted distributions count every observation. Photographer-balanced distributions give each photographer equal total weight within the displayed metric and context. Pages also show the largest contributor share so concentrated samples remain apparent.
Settings buckets and completeness
ISO uses bounded ranges; aperture is rounded to one decimal place; shutter speed uses one-third-stop logarithmic buckets; focal lengths use the nearest millimetre. Field coverage is counted separately, and missing values are not treated as zero.
Indexing quality
A URL must belong to the reviewed publication allowlist and pass its runtime sample, photographer, concentration, completeness, and content gates. Passing a numerical threshold does not automatically create a page or imply that every possible camera, lens, and genre combination should be indexed.
Version lifecycle and limitations
Imports build a separate version and activate it only after validation. Normalization, classification, quality aggregates, and settings distributions are rebuilt transactionally. Results reflect the current contributing sample, may change between versions, and cannot establish causation or universal practice.