How the indices are built
This page is generated from the internal methodology specification, so the published method and the one the code implements cannot drift apart. Every weight, cap and threshold below was fixed before any data existed — deriving weights from data once it accrues produces a series whose history changes each time you refit it, which is unusable for anyone modelling against it.
Version 0.1, pre-registration draft. Four parameters remain open: the discount rate used in Realized Farmgate Spread, confirmation of the duration caps, the regional post-stratification weights, and the publishable cut list. They will be fixed before Wave 1 locks and this page will record the date they were.
Principles
Pre-registered, not fitted. Weights, transforms and caps are fixed in this document before data exists. Deriving weights from the data once it accrues — PCA, regression against arrivals — produces a series whose history changes every time you refit it. That is unsellable to a quant.
Every index decomposes. A subscriber can always see the raw components in natural units. The index is a convenience, never the only view.
Polarity is stated, never coloured. Payment Stress rising is bad for farmers and informative for traders. We state direction and let the subscriber decide what is good.
Suppression over estimation. Where sample is too thin, we publish nothing and say why. Consistent with how the external-data side already handles absent sources.
The Wave 1 base problem
Two facts collide. Wave 1 is our first observation, so any base-100 index makes Wave 1 informationless. And cocoa is intensely seasonal — a main-crop October with freshly funded LBCs is structurally different from a mid-crop July, so a base fixed in one month contaminates every comparison until we have years of history for seasonal adjustment.
Decision: the two headline indices are bounded 0–100 constructions, not base-100 indices. They are interpretable from the first release with no base period at all, in the way PMI is. A reading of 34 on Purchasing Capacity means something on day one.
Held Stock cannot escape being a level, and Realized Farmgate Spread is a ratio that needs no base. Handled separately below.
No seasonal adjustment before 36 monthly observations. Until then every release carries its crop-calendar phase label, and month-on-month comparisons carry an explicit warning that seasonality is not removed. Year-on-year becomes available from Wave 13. Saying this plainly is better than being asked.
Purchasing Capacity Index (PCI)
The most important measure in the product. Purchasing clerk funding status is a leading indicator of official arrivals that no one publishes, and it is the single strongest argument for the panel existing.
Polarity: higher = more capacity to buy. Falling PCI is deteriorating.
Respondents: purchasing clerks (primary), depot/LBC (secondary).
| Component | Source question | Transform | Weight |
|---|---|---|---|
| Share fully funded | PC funding status | p × 100 | 0.35 |
| Funding recency | Days since funds last received | 100 − scaled(days, CAP=21) | 0.25 |
| Turn-away rate | Bags refused because PC could not pay | 100 − scaled(share_of_offered, CAP=0.25) | 0.25 |
| Evacuation currency | Days since last evacuation | 100 − scaled(days, CAP=28) | 0.15 |
PCI = Σ (weight × component_score)
Weights are judgment, weighted toward funding status because it is the most direct measurement and the least noisy. Publish the four components in natural units in every release.
Minimum n: 30 purchasing clerks per published cell.
Payment Stress Index (PSI)
Polarity: higher = more stress. Rising PSI means farmers are being paid worse or slower.
Respondents: farmers.
| Component | Transform | Weight |
|---|---|---|
| Share of deliveries not paid in full | p × 100 | 0.30 |
| Mean days delivery→payment (settled) | scaled(days, CAP=30) | 0.25 |
| Mean days currently outstanding (unsettled) | scaled(days, CAP=45) | 0.25 |
| Chit / credit-note usage | p × 100 | 0.20 |
Two duration components with different caps: a settled 30-day payment and a 30-day-and-counting outstanding balance are different conditions, and the second is worse.
Farmers reporting no delivery since last wave are excluded from PSI and counted separately. Their exclusion is itself a signal — a rising share of non-delivering farmers alongside rising Held Stock is the pattern the product exists to catch. Publish that share as PSI_base_share.
Minimum n: 50 farmers reporting a delivery per published cell.
Realized Farmgate Spread (RFS)
Commercially the most interesting measure, and the fiddliest. Published as a percentage, not an index.
The naive version — agreed price versus official price — misses the point. A price agreed but paid in 40 days is not the official price. Two variants, both published:
RFS-nominal — price discovery.
realized_nominal = (agreed_price + premiums − deductions) / kg RFS_nominal = (realized_nominal − official_price) / official_price × 100
RFS-realized — what the farmer actually got, time-adjusted.
PV = cash_received + outstanding / (1 + r)^(days_outstanding / 365) realized_cash = PV / kg_delivered RFS_realized = (realized_cash − official_price) / official_price × 100
r is a pre-registered annual discount rate. Decision needed: use the Ghanaian commercial lending rate, refreshed annually, or a fixed stated rate. I recommend the lending rate — it reflects the farmer's actual cost of not having the money — with the rate used printed in every release.
Units: Ghana quotes per 64 kg bag. Store as given, convert to per-kg and to USD per tonne at the release-date FX rate. Publish all three. Never silently convert.
Cross-border variant (from Wave 1 in Côte d'Ivoire) compares realized Ghana against realized Côte d'Ivoire on an FX-adjusted USD/t basis. This is the smuggling-arbitrage indicator and probably the single most saleable number in the platform.
Minimum n: 50 priced transactions per cell.
Held Stock Index (HSI)
The weakest of the four in year one, and we should say so publicly rather than have a subscriber discover it. It is a level with strong seasonality and no base period.
Primary publication is absolute, not indexed:
- Mean bags held unsold per farmer reporting stock
- Share of farmers holding any unsold stock
- Mean bags in PC shed
- Mean depot inventory in tonnes
- Mean days since last evacuation
Secondary is a composite indexed to Wave 1 = 100, labelled explicitly as not seasonally meaningful until Year 2. Weights: farmer-held 0.4, PC shed 0.35, depot 0.25 — reflecting where stock accumulation is most diagnostic of a marketing-system failure rather than normal working inventory.
Held stock must be published alongside the stated reason distribution (waiting for price, buyer unavailable, PC lacks money, transport, quality). The reason is what separates a functioning market holding inventory from a seizing one.
Sampling weights
Decision: post-stratify to regional shares of official cocoa purchases, using COCOBOD regional purchase data, with weights fixed for the season and revised at the annual review.
Unweighted means would let a region with more enumerators dominate. Volume weighting at the respondent level is better in principle but needs a volume denominator we will not have in Wave 1 — revisit at the Year 1 review.
Within a region, respondents are equally weighted, except for price and volume measures where farmer observations are weighted by bags delivered in the reference period.
Guardrail: no single respondent may carry more than 20% of the weight in any published cell. If one does, the cell is suppressed.
Missing data
Distinguish two failures that are usually conflated.
Item non-response — respondent interviewed, question unanswered. Compute on available responses. Never impute. Publish the base n for every figure.
Unit non-response — respondent missed the wave entirely. This is panel attrition and it is dangerous, because attrition is rarely random: a purchasing clerk who has stopped operating is exactly the observation you most need.
Decision: enumerators must record a non-response reason — unreachable, refused, ceased trading, relocated, deceased. Ceased trading is a data point, not a gap, and feeds Purchasing Capacity directly rather than vanishing.
Matched-sample rule: any wave-on-wave change figure is computed only on respondents present in both waves. Level figures use the full available sample. Both n's are published. Without this, panel churn produces movement that looks like market movement.
Publication thresholds
| n in cell | Action |
|---|---|
| ≥ 30 (≥ 50 for PSI/RFS) | Publish |
| 15–29 | Publish with a low-sample flag; excluded from headline |
| < 15 | Suppress; show "insufficient sample", not a blank |
| Any respondent > 20% of weight | Suppress |
Applies to every geographic and respondent-class cut. Decide the publishable cut list before Wave 1, because 400 respondents across five regions and five classes leaves several cells in single digits by construction.
Hauliers: narrative only. Decided 27 August 2026. At n=10 the panel cannot support an index component, so haulier output is published as written commentary in the release note and is excluded from Physical Flow.
The data is nonetheless collected in structured, quantified form — trip counts, tonnage, origin district, destination type, delay hours. Free text cannot be indexed retrospectively, so collecting narrative would permanently foreclose the option. If the haulier panel later reaches n>=40, Physical Flow gains a component with its history already in place.
Revision and vintage policy
The stated policy.
1. A wave locks at publication. Locked values are never edited. 2. A correction creates a new vintage: prior value, corrected value, reason, analyst, timestamp. Both remain retrievable. 3. Real-time vintages are preserved and published. What did the index say on the day it was released, not what does it say now after revision. Quant subscribers pay for this and almost no alternative-data vendor offers it. 4. Weight re-stratification at annual review triggers a full historical recomputation, published as a new vintage with both series available in parallel for one year. 5. A cap or transform change follows the same path and may happen once.
What we will not claim
Worth writing down now, because the pressure to overclaim arrives with the first sales call.
- Not a forecast. These measure current physical conditions. Any lead-lag relationship to arrivals or price is a hypothesis for subscribers to test, and we publish the data to let them, rather than asserting it.
- Not representative of all Ghanaian cocoa. A 400-respondent purposive panel in five regions. State the coverage, state what is excluded.
- Not seasonally adjusted before Year 3.
- Not a substitute for official statistics. A different measurement of a different thing, at a different speed.
Questions this page should answer
If it does not answer yours, ask us. Methodology questions from prospective subscribers have changed this document more than once.