Methodology
How PhotoDataLab analyzes photography metadata
PhotoDataLab turns selected metadata from public portfolios of professional photographers into versioned observational datasets. Sample size, contributor diversity, field coverage, and uncertainty stay visible beside every result.
Observations, not universal advice
Results use language such as “observed” and “most frequently observed.” A common value is not automatically an optimal value, a recommendation, or a measure of market popularity.
Scope and exclusions
Dataset version 2 contains 166,036 photo observations from 1,063 professional photographers across 5,537 public albums. Photo files are not copied or embedded. Unnecessary personal information and GPS coordinates are excluded from the public dataset.
“Professional photographer” describes the source population. It does not establish that every observation came from a paid assignment, or that this dataset represents the entire profession, a country, a market segment or every photographic specialty.
Camera and lens identities
Raw EXIF make, model, and lens strings are cleaned by deterministic versioned normalizers. Alias records preserve how a value appeared while mapping proven formatting and manufacturer-prefix variants to one canonical identity. Ambiguous generic strings are not presented as specific products.
Genre classification
Genres are inferred from explicit keywords in public album names by classifier
album-keywords-v1. This is a reproducible heuristic, not a human-verified label for every
photo.
Contributor balance
Photo-weighted distributions count every observation. For a photographer-balanced distribution, each photographer's observations in one displayed metric and context are divided among its buckets so that photographer contributes a total weight of one. A bucket's fractions are summed across photographers and divided by the total weight across all buckets. The displayed percentage therefore gives every professional photographer equal overall influence, regardless of portfolio size. Photographers without a value for that metric are excluded from its distribution. Pages also show the largest contributor share so concentrated samples remain apparent.
Settings buckets and completeness
ISO uses bounded ranges; aperture is rounded to one decimal place; shutter speed uses one-third-stop logarithmic buckets; focal lengths use the nearest millimetre. Field coverage is counted separately, and missing values are not treated as zero.
Indexing quality
A URL must belong to the reviewed publication allowlist and pass its runtime sample, photographer, concentration, completeness, and content gates. Passing a numerical threshold does not automatically create a page or imply that every possible camera, lens, and genre combination should be indexed.
Public source references
Some observations have a verified public custom-domain trail. Source pages show public album and contributor-website links only when their URL can be constructed from that verified domain. PhotoDataLab does not guess missing links, publish internal identifiers, reproduce images or imply an affiliation with the linked publisher. Observations without a verified URL can still contribute to aggregate statistics.
Version lifecycle and limitations
Imports build a separate version and activate it only after validation. Normalization, classification, quality aggregates, and settings distributions are rebuilt transactionally. Results reflect the current contributing sample, may change between versions, and cannot establish causation or universal practice.