Finding 01
Variability in units sold for top-performing SKUs
Even a top-selling body wash showed a wide range in demand over the period studied.
What Pavo used: Sales history, alongside other signals, to help represent changing demand.
Case study
A house of consumer brands
Amazon-native · multi-brand consumer group
Pavo built a demand model that learns across products and uses advertising and pricing inputs. In reported backtests, 2-month-ahead accuracy rose from 28% to 62%.
Reported backtest accuracy. The source deck does not define the accuracy formula or weighting.
The problem
For an Amazon brand, forecasting demand means deciding how much of each product to stock. Underestimate demand and sales can be lost to stockouts. Overestimate it and cash is tied up in unsold inventory.
The reported starting point was 19% accuracy 1 month ahead, 28% at 2 months, and 13% at 3 months. Pavo set out to improve forecasts across all 3 planning horizons.
Pavo started with 2 brands. It examined their sales and product data, tested new inputs, and compared models using backtests: predictions evaluated against historical outcomes.
The change
A SKU is an individual product or variant. Forecasting each SKU in isolation limits it to its own history. Pavo trained across SKUs so products with little history could benefit from patterns learned elsewhere in the catalog. Each product still gets its own forecast.
Product by product
One product’s sales history informs its own forecast.
Across products · Cross-SKU
Shared learning, with a separate forecast for each product.
Shared forecasting model
Also uses ratings, product categories
and seasonality features
The analysis
Pavo examined how demand changed over time and across products. Each finding below connects an observation in the data to an input or training choice: recent demand, seasonality, reviews, or learning across SKUs.
Finding 01
Even a top-selling body wash showed a wide range in demand over the period studied.
What Pavo used: Sales history, alongside other signals, to help represent changing demand.
Finding 02
The selected windows show similar late-year rises from different starting points: roughly 113 to 149 units in 2022, and 45 to 78 in 2023. Pavo treated this as exploratory evidence of a recurring pattern.
Monthly moving-average units
What Pavo used: Seasonality features to represent the rise, alongside recent-demand features to represent the sales level.
Finding 03
The share of 5-star reviews was more closely associated with units sold than the shares of lower ratings.
Correlation with units sold
What Pavo used: Rating features as additional inputs to the forecasting model.
Finding 04
Two similar body washes had positively correlated sales values. A body wash and a hair treatment showed the opposite relationship.
Sales value (GMV) correlation
Body wash + body wash
Listing similarity 0.96 · GMV correlation +0.42
Body wash + hair treatment
Listing similarity 0.29 · GMV correlation −0.46
What Pavo used: Cross-product learning and product features to give each forecast more context.
The approach
Pavo moved from the group's operating context to a tested forecasting system in seven steps.
Absorbed the system knowledge
Connected to Amazon sales, advertising, pricing, and ratings data for both brands, and took the group's current forecast, 61%, 31%, and 40% at 1, 2, and 3 months on the personal-care brand, as the bar to beat.
Diagnosed the dynamics
Traced the variability, seasonality, ratings correlation, and cross-SKU correlation above, and confirmed that some SKUs had years of history while others had months.
Encoded the findings as features
Built 8 feature families: seasonality; Amazon sales velocity; category and purchase frequency; holidays, week, and month of year; lagged and delta features; review scores, text, and aspects; category-tree rank; and listing-aspect keywords for quality, accessibility, and eco claims.
Trained one model across all SKUs
Instead of fitting one model per SKU, Pavo trained one model on every SKU jointly, so a sparse SKU could borrow seasonal shape and price response from similar SKUs.
Found the input that mattered most
Next month's ad spend and average selling price are largely the group's own decisions, so Pavo tested supplying them to the forecast. Comparing them withheld and supplied produced the largest single effect of the engagement.
Compared 7 architectures per horizon
Pavo scored a Prophet-based model, two combined MLPs, two Autoformer variants, nLinear, and dLinear on the same backtest at 1, 2, and 3 months. Most alternatives deteriorated materially by 3 months; only the Prophet-based model held across all 3 horizons.
Built the custom model that won
Pavo-Prophet carried the full feature set, cross-SKU training, ad spend, and average selling price into a Prophet-based model tuned for this data. Pavo scored it on both brands at all 3 horizons.
The key insight
Future ad spend and average selling price (ASP) added information that past sales alone could not provide. With those inputs supplied, the personal-care brand's 2-month accuracy rose from 28% to 62%. The charts show the effect at every horizon for both brands.
Backtest accuracy · higher is better
+16 percentage points
+34 percentage points
+21 percentage points
Backtest accuracy · higher is better
+33 percentage points
+21 percentage points
-9 percentage points
The source deck describes ad spend and ASP as known inputs. It does not document whether each value was available at the historical forecast cutoff or realized later.
Features & training design
Sales history shows what sold. Pavo added features describing recent momentum, seasonal timing, and product context. It also trained across SKUs with different amounts of history, changing both what the model sees and which products it learns from.
01 / Enrich the inputs
How is demand moving?
Where are we in the sales cycle?
What kind of product is this?
How is the product described and rated?
02 / Change how the model learns
Pavo trained across the full SKU catalog. Products with months of data learned alongside products with years of history, while product context kept each forecast specific to its SKU.
Input
Observed before the cutoff
Sales velocity · lags and deltas · reviews · category context
Supplied for the forecast window
Calendar and holidays · supplied ad spend · supplied average selling price
Shared cross-SKU model
Learns from every SKU jointly
Combines shared patterns with each product's history and context
Output
Forecasts for each SKU
Trying parallel approaches
Pavo tested Autoformer, dLinear, nLinear, and a custom Prophet-based model, alongside MLP variants. The four families below handle trends, recurring patterns, and changes in sales level differently. The next comparison shows how the tested variants performed at each forecast horizon.
01 · Learn repeating patterns
Autoformer progressively separates longer-term trends from seasonal fluctuations inside its network. It uses autocorrelation to find repeating patterns across time.
Personal-care brand baseline model
Repeated inside the network
02 · Split, project, add
dLinear separates sales history into a moving-average trend and a remainder. It projects each part forward with a separate linear layer, then adds the results.
Personal-care brand baseline model
03 · Adjust around the latest value
nLinear subtracts the most recent observation from the sales history before applying one linear layer. It adds that observation back to express the forecast at the original sales level.
Personal-care brand baseline model
04 · Combine interpretable components
Prophet combines trend, seasonal and holiday effects, with support for additional inputs. The comparison in this case study uses Pavo’s custom Prophet-based variant.
Personal-care brand baseline model
Simplified additive formulation
Offline evaluation: comparing different models
Strong results next month did not guarantee strong results further out. The larger MLP scored 72% at 1 month and 9% at 3 months. Pavo-Prophet, the custom Prophet-based model, scored 80%, 62%, and 50%, the highest reported accuracy at each horizon.
Backtest accuracy · higher is better
Backtest accuracy · higher is better
Backtest accuracy · higher is better
The result
Pavo compared each brand's baseline forecasting model with the final Pavo-Prophet model at 1, 2, and 3 months ahead.
Backtest accuracy · higher is better
+59 percentage points
+34 percentage points
+37 percentage points
Backtest accuracy · higher is better
+34 percentage points
+15 percentage points
+6 percentage points
Teams using Pavo to move the metrics that matter, in production, at scale.