Field notes

SKU-Level Carbon Footprints: From Hero Products to Your Whole Catalog

Scaling from a few hero-product LCAs to SKU-level footprints across your catalog: what breaks, what to standardize, and how to keep numbers defensible.

Field notes4 min readUpdated July 22, 2026

Most product-carbon programs start the same way: a consultant-built LCA for three to five "hero products," namely the bestseller, the new launch, and the one a big retail customer asked about. Those studies are careful, expensive, and finished. Then a buyer, an auditor, or CSRD asks about the other 2,000 SKUs, and the approach that produced five good numbers turns out not to scale to two thousand.

The gap between a hero-product LCA and a catalog-wide SKU program isn't effort. It's method. Here is what actually changes, and what to standardize before you scale, so the numbers you produce at volume are still numbers you can defend.

Why hero-product methods don't scale

A hand-built LCA absorbs inconsistency invisibly. When one analyst models five products over eight weeks, they make dozens of judgment calls, such as which emission factor database, which allocation rule for shared processes, and how to treat transport legs with missing data. They make them consistently because one brain holds them all.

Spread those same calls across 2,000 SKUs, multiple analysts, and a year of updates, and the judgment calls drift. Two near-identical products end up with different factor versions. A packaging change gets reflected in one SKU but not its siblings. The result is a catalog where differences between products reflect modeling noise as much as real emissions, which is precisely what a buyer comparing your SKUs, or an assurer sampling them, will find.

The three things to standardize first

1. A component library, not per-product models. Hero-product LCAs model each product from scratch. At catalog scale you invert it: model materials, processes, packaging formats, and transport legs once, as shared components with pinned emission factors, then compose SKUs from the library. A fabric, a trim, or a carton modeled once flows to every SKU that uses it, and an update to that component propagates everywhere at once instead of creating version skew.

2. One factor-selection policy, written down. Decide once: which databases, in what order of preference, which vintage, and when a supplier-specific primary factor overrides a secondary one. The policy matters more than the choices, because an emission factor is traceable only if you can name its source and version, and at 2,000 SKUs "the analyst picked something reasonable" is not a policy an assurer can test.

3. Allocation rules per shared process. Dye houses, shared freight, co-packed lines: wherever one activity serves many SKUs, the split rule, whether mass, unit, or economic value, must be fixed per process and applied uniformly. Inconsistent allocation is the most common way two similar SKUs end up with implausibly different footprints.

Scale in tiers, not all at once

No one models 2,000 SKUs at hero-product depth, and no one needs to. A workable rollout looks like this:

  • Tier 1, the SKUs that get scrutiny. Top sellers by volume, plus anything a customer or regulator has asked about. Full model, supplier-specific primary data where it exists, reviewed individually. Typically 5 to 10% of the catalog and often 40% or more of unit volume.
  • Tier 2, the body of the catalog. Composed from the component library with secondary factors, flagged wherever a primary-data gap moves the number materially. Modeled in batches, sampled for review.
  • Tier 3, the long tail. Low-volume and discontinued-soon items get screening-level estimates, clearly labeled as such, queued for upgrade only if something promotes them: a reorder, a disclosure request, a hotspot flag.

The tier labels themselves are worth disclosing internally. A screening estimate presented as a measured footprint is the kind of overstatement that falls apart under assurance questioning. A screening estimate labeled as one is a legitimate, defensible stage of a maturing program.

Keep the evidence chain as you scale

Here's the trap: scaling pressure pushes teams toward whatever produces numbers fastest, and the first thing dropped is the paper trail. The hero-product LCA had an appendix citing every source. The catalog rollout has a spreadsheet of outputs and no record of which factor version, which supplier document, or which allocation rule produced each one.

That trade-off is invisible right up until someone challenges a number: an assurer sampling SKU #1,347, a retailer comparing your footprint to a competitor's, or an ESRS E1 disclosure that has to survive limited assurance. Recomputing the answer from scratch a year later is usually impossible, because the analyst has moved on and the factor database has a new vintage.

The fix is structural, not heroic. Every SKU-level result should carry its lineage, meaning activity data in, factor and version applied, allocation rule, and result out, captured at calculation time rather than reconstructed later. This is the core argument of the evidence-ledger approach, and it's the problem CarbonSKU is built around: each number bound to its versioned source documents, so SKU #1,347's footprint is as replayable as the hero product's. However you implement it, the test is simple. Pick a random SKU and ask why the number is what it is. If the answer requires archaeology, the program has scaled past its evidence.

What "done" looks like

A catalog-scale PCF program is working when any SKU's number traces to its components, factors, and rules in minutes; when two similar SKUs differ only for reasons you can name; when a component update propagates to every affected SKU with a version history; and when the tier label on every number states its depth plainly. That's the difference between having 2,000 numbers and having 2,000 numbers you'd let an auditor pick from at random.