Stop the leak: 10 AI skills for what your suppliers charge you

supplier-scorecard

grade the suppliers you already use, in numbers

How the two work together

Claude thinks it through. Paste the Claude prompt into Claude Code, or drop the folder into your skills folder. Claude does the judgement: what to look for, what is worth doing, what is right.

Codex gets it done. At the hand-off point Claude runs Codex on your machine with one command and passes it the Codex prompt. Codex does the mechanical part and hands the result back. Claude checks it before you see it.

No API key to set up: Claude calls the Codex you already have installed. If Codex is not installed, Claude does that half itself and tells you.

Prompt for Claude

---
name: supplier-scorecard
description: Grades the suppliers you already buy from on lateness, short deliveries, substitutions, price rises and credits, using your own delivery notes and invoices, so a renewal or price-review conversation has figures in it instead of opinions.
---

# What your suppliers actually delivered, in numbers you can put on the table

You give this your last three months of orders, delivery notes, invoices and credit notes for one supplier. You get back a one-page scorecard: how often they turned up when they said, how often the order was complete, what they swapped without telling you, how much their prices moved in pounds, and how long they took to pay a credit. It measures the supplier you already have. It is not for comparing quotes from new ones.

## What it does

1. **Fix the window at the last 13 full weeks and pull every piece of paper for it.** For each delivery you need four things: what you ordered (the order confirmation, the email, or your written note of a phone order), what the delivery note says arrived, what the invoice charged, and any credit note raised. If a delivery is missing one of the four, leave it out of the scoring and count it on a separate "paperwork missing" line. A supplier with 30 deliveries and 9 incomplete sets has already told you something about how they run your account.

2. **Build one row per delivery before you calculate a single percentage.** Use these columns and no others: delivery date, slot you were promised, time it actually arrived, order lines, lines delivered in full, lines short, lines substituted, lines you rejected, order value, invoice value, credit claimed, credit received, date the credit landed, document references. Every figure in every cell carries the delivery note or invoice number it came from. If a number has no document number beside it, it does not go on the scorecard.

3. **Score on time and in full separately, then combine them.** Wikipedia's article on DIFOT gives the sum as "OTIF (%) = number of OTIF deliveries ÷ total number of deliveries × 100", and stresses that the point of it is to judge delivery "from the point of view of the customer". So record two dates on every row: the slot you asked for and the slot they later confirmed. The same article notes research by Janet Godsell finding that suppliers often hit their targets against their own promised dates and miss badly against the dates the customer actually asked for. Report both figures. The gap between them is usually the argument.

4. **Count fill rate by order line, not by whole orders.** Wikipedia's DIFOT article notes the calculation may be done "by orders or order lines rather than deliveries", and for a kitchen the line is the honest unit: one missing box of gloves should not wipe out a 40-line delivery, and it should not disappear either. Report two numbers: lines delivered in full as a percentage of lines ordered, and the percentage of deliveries that were 100% complete with nothing short, nothing swapped. The second number is usually far lower than the supplier's own figure and it is the one your chefs feel.

5. **List every substitution by name, with both prices.** GoodSource's guidance for foodservice distribution says substitution standards have to be agreed in advance, and gives the case of a distributor replacing "a 10-pound case of chicken breasts with two 5-pound cases at the same per-pound price". For each swap record the item ordered, the item delivered, the unit price of each, and whether you were told before the lorry left the depot. Count any substitution you were not warned about as a failed line, even if the price was identical, because an unannounced swap is also a menu description and an allergen problem.

6. **Work out price drift in pounds against the volumes you actually bought.** Take the agreed price list, or the first invoice in the window if there is no list, as the baseline for every product. Compare each later invoice line against it. GoodSource treats a running variance of "3-5% variance" as the consistency band for foodservice pricing, so flag every product that moved more than 5% and every one that moved more than once. Then multiply each per-unit move by the units you genuinely bought in the 13 weeks and add it up. An average percentage is easy to shrug off; "this cost me £1,840 over the quarter, and £1,100 of it is these four products" is not.

7. **Put a clock on every credit.** Count credits claimed, credits actually received, pounds still outstanding, and the age in days of the oldest unpaid claim. Then count the shorts on delivery notes that never became a claim at all, because those are money you have already written off without noticing. Report the average days from claim to the credit appearing on a statement. A supplier who is 96% on time but takes six weeks to settle a £40 credit is charging you in admin what they saved you in price.

8. **Check the temperature and rejection record against the legal line, not a house rule.** GOV.UK's food hygiene guidance for businesses states: "Chilled food must be kept at 8°C or below. This is a legal requirement in England, Wales and Northern Ireland", and adds "Set your fridge to 5°C or below to make sure food stays cold enough, even if the temperature changes." Count the deliveries that arrived above 8°C, whether each was rejected, and whether the rejection was written down at the time. Do the same for frozen loads that arrived soft. These rows belong in your food safety records as well as on the scorecard.

9. **Weight the five categories, grade the supplier and write the single page you take into the meeting.** Use delivery reliability 35%, order accuracy and substitutions 25%, price 20%, credits and admin 10%, temperature and condition 10%, each scored out of 100. EvaluationsHub's scorecard guidance advises setting "guardrails so one KPI cannot dominate the total score", so no category can earn above its weight and a good price cannot buy back a bad delivery record. Grade the weighted total as A at 85 or above, B at 70 to 84, C at 55 to 69 and D below 55; the asgard-ai-platform supplier scorecard skill pairs its C grade with an "improvement plan required within 90 days", which is the right ask for a renewal rather than walking away. GoodSource recommends monthly scorecards land "within 48 hours of the period closing", so send it within five working days, put the three specific asks and the next review date at the bottom, and keep the workings attached.

## Then it checks

1. Every percentage, count and pound figure on the finished page traces to a listed delivery note, invoice or credit note number, and the page carries that evidence list.
2. On-time percentage and in-full percentage appear as two separate figures as well as a combined OTIF figure, each showing the number of deliveries behind it, and the on-time figure is shown twice: against the slot you requested and against the slot the supplier confirmed.
3. Every substitution inside the window is listed on its own line with the item ordered, the item delivered, both unit prices, and a yes or no for whether you were told in advance.
4. Price drift is reported in pounds against real purchased volumes as well as in percentages, and the individual line movements add up to the stated total.
5. The credits section shows claimed count, received count, pounds outstanding, average days to settle, and the age of the oldest unpaid claim.
6. The page states the exact date range and the number of deliveries scored, and where that is fewer than 20 deliveries or under 8 weeks it says on its face that the sample is too small to grade and shows every figure as indicative only.

Any check fails: name it, redo that step once. Failed twice: say what is wrong and stop.

## Rules
- Public information only.
- Never invent a fact, a number or a quote.
- Never fill a gap in the paperwork with an estimate, a typical figure or a rounded guess. A scorecard with three honest holes in it survives being challenged by the account manager; one invented average destroys the whole page the moment they check that line, and you lose the argument on the points you were right about.
- Never average a serious single failure into a comfortable quarterly percentage. One load of chicken at 12°C, one unannounced substitution that broke an allergen description, or one delivery that never arrived on a Saturday service must appear as its own named incident with its date, not as a rounding error inside a 94% score.
- Never convert a supplier's behaviour into a claim, a deduction or a threat to terminate. This output covers food safety records and money you say you are owed, so it is a working document prepared for your solicitor, your accountant and your environmental health officer to check before you act on it, and it does not pretend to be their advice.

## Built from
- Wikipedia, "DIFOT", https://en.wikipedia.org/wiki/DIFOT, last edited 20 July 2026: gave the OTIF formula used in step 3, the choice between counting deliveries, orders or order lines used in step 4, and the Godsell finding that suppliers pass against their own promised dates and fail against customer-requested ones.
- GoodSource, "Vendor Performance Evaluation Metrics for Wholesale Food Distribution Partnerships", https://goodsource.com/trends-and-insights/vendor-performance-evaluation-metrics-for-wholesale-food-distribution-partnerships/, 22 March 2026: gave the foodservice substitution example in step 5, the 3 to 5% price variance band in step 6, and the rule that a scorecard is worthless if it arrives long after the period closed.
- EvaluationsHub, "Supplier Scorecard Best Practices: KPIs, Weighting, Cadence", https://evaluationshub.com/supplier-scorecard-best-practices-kpis-weighting-cadence/, 4 March 2026: gave the weighting guardrail in step 9 that stops one measure dominating the total, and the principle of defining each measure in auditable terms before scoring anything.
- GOV.UK (Food Standards Agency guidance), "Food hygiene for businesses: Chilling and freezing", https://www.gov.uk/food-hygiene-businesses/chilling-and-freezing, no publication date shown, read 13 September 2026: gave the 8°C legal limit and the 5°C working target used as the pass or fail line for delivery temperatures in step 8.
- asgard-ai-platform, "mfg-supplier-scorecard/SKILL.md", https://github.com/asgard-ai-platform/skills/blob/main/mfg-supplier-scorecard/SKILL.md, 228 stars read from api.github.com, no publication date shown, read 13 September 2026: gave the four-dimension weighted structure, the A to D grade bands and the 90-day improvement plan attached to a C grade in step 9.

Prompt for Codex

# supplier-scorecard

## You are given
A folder from a UK hospitality business covering one supplier and the last thirteen full weeks: the orders (order confirmations, emails, or written notes of phone orders), the delivery notes, the invoices, the credit notes, the statements, the agreed price list if one exists, and the temperature and rejection records taken at the door. A text file gives the supplier name, the account number, the window start and end dates, and the delivery slot requested for each order where that differs from the slot the supplier confirmed. Assume several deliveries are missing one of the four documents, and that some substitutions were never written down anywhere but the delivery note.

## Produce
Write into an `output/` folder next to the inputs:

1. `deliveries.csv` with these columns in this order and no others: `Delivery date`, `Slot requested`, `Slot confirmed by the supplier`, `Time it actually arrived`, `Order lines`, `Lines delivered in full`, `Lines short`, `Lines substituted`, `Lines rejected`, `Order value net of VAT (GBP)`, `Invoice value net of VAT (GBP)`, `Credit claimed (GBP)`, `Credit received (GBP)`, `Date the credit landed`, `Order reference`, `Delivery note number`, `Invoice number`, `Credit note number`. Every figure in every cell carries the document number it came from in the reference columns.
2. `paperwork-missing.csv` with columns: `Delivery date`, `Order held` (yes/no), `Delivery note held` (yes/no), `Invoice held` (yes/no), `Credit note expected` (yes/no/none), `Which document is missing`. These deliveries are excluded from every percentage and counted here only.
3. `otif.csv` with columns: `Measure`, `Numerator`, `Denominator`, `Percentage`, `What it counts`. Rows in this order: `On time against the slot requested`, `On time against the slot the supplier confirmed`, `In full by order line`, `Deliveries 100% complete`, `OTIF against the slot requested`, `OTIF against the slot confirmed`.
4. `substitutions.csv` with columns: `Delivery date`, `Item ordered`, `Code ordered`, `Item delivered`, `Code delivered`, `Unit price ordered net of VAT (GBP)`, `Unit price delivered net of VAT (GBP)`, `Difference (GBP)`, `Told in advance` (yes/no), `Delivery note number`. Every substitution in the window appears on its own line.
5. `price-drift.csv` with columns: `Supplier product code`, `Description`, `Baseline price net of VAT (GBP)`, `Baseline source` (agreed price list and date, or first invoice in the window and its number), `Latest price net of VAT (GBP)`, `Latest invoice number`, `Latest invoice date`, `Movement per unit (GBP)`, `Movement %`, `Separate moves in the window`, `Units bought in the window`, `Cost of the movement (GBP)`. Sorted by `Cost of the movement (GBP)`, highest first, with the column total at the foot.
6. `credits-clock.csv` with columns: `Claim date`, `Delivery note number`, `Original invoice number`, `Claimed (GBP)`, `Received (GBP)`, `Outstanding (GBP)`, `Credit note number`, `Date it landed on a statement`, `Days from claim to credit`, `Status`. Below the table: credits claimed, credits received, pounds outstanding, average days to settle, age in days of the oldest unpaid claim, and the count of shorts written on delivery notes that never became a claim.
7. `temperature-and-rejections.csv` with columns: `Delivery date`, `Time`, `Product`, `Product code`, `Reading (degrees C)`, `Chilled or frozen`, `Above 8 degrees C` (yes/no), `Frozen load arrived soft` (yes/no), `Rejected` (yes/no), `Written down at the time` (yes/no), `Source file`.
8. `incidents.md` - one numbered entry per serious single failure, each with its date, the document numbers, and what happened, written as facts. A load above 8 degrees C, an unannounced substitution that changed a menu description or an allergen, and a delivery that never arrived each get their own entry. None of these is averaged into a percentage anywhere.
9. `scores.csv` with columns: `Category`, `Weight %`, `Score out of 100`, `Weighted score`, `Figures behind it`. Rows in this order and with these weights: `Delivery reliability` 35, `Order accuracy and substitutions` 25, `Price` 20, `Credits and admin` 10, `Temperature and condition` 10. Final rows: `WEIGHTED TOTAL` and `GRADE`. Grade A at 85 or above, B at 70 to 84, C at 55 to 69, D below 55.
10. `scorecard.md` - one printed page: the supplier name and account number, the exact date range, the number of deliveries scored, the five categories with their weighted scores, the grade, the three headline figures (OTIF both ways, cost of price movement, pounds of credit outstanding), the incident list, and the next review date. The three specific asks and the workings list go at the bottom.
11. `README.md` - the files read, the exact date range, the number of deliveries scored, the number excluded for missing paperwork, and the evidence list of every document number used.

## Rules
- The window is the last thirteen full weeks. A delivery missing any of the four documents is excluded from scoring and counted in `paperwork-missing.csv`.
- Every percentage, count and pound figure must trace to a listed delivery note, invoice or credit note number. A number with no document number beside it does not go on the scorecard.
- Report on time and in full as two separate figures as well as combined, and report on time twice: against the slot requested and against the slot the supplier confirmed.
- Count fill rate by order line, not by whole orders, and report the count of deliveries that were 100 per cent complete alongside it.
- Count any substitution the business was not warned about in advance as a failed line, even where the price was identical.
- Flag every product that moved more than 5 per cent and every product that moved more than once. Price drift is reported in pounds against the units actually bought as well as in percentages, and the line movements must add up to the stated total.
- Never fill a gap in the paperwork with an estimate, a typical figure or a rounded guess.
- Never average a serious single failure into a quarterly percentage. It appears as its own named incident in `incidents.md` with its date.
- No category may score above its own weight, so a good price cannot buy back a bad delivery record.
- Where fewer than 20 deliveries were scored, or the window is under 8 weeks, write `Sample too small to grade. Every figure below is indicative only.` at the top of `scorecard.md` and in `README.md`, and still show the figures.
- All values go in net of VAT and are labelled as such.
- Never convert a figure into a claim, a deduction, a threat to terminate or a recommendation to leave. State the numbers.
- Round only at the point of display, to two decimal places.
- Use British English, £, and DD Month YYYY dates. No em dashes.
- The output is a working document prepared for the owner's solicitor, accountant and environmental health officer to check. It covers food safety records and money the business says it is owed. Write no sentence that presents it as their advice.

## Return
A list of the files written with their absolute paths, the number of deliveries scored and the number excluded, the on-time percentage against both slots, the in-full percentage by line, the cost of price movement in pounds, the pounds of credit outstanding and the age of the oldest unpaid claim, the weighted total and the grade, and every entry in `incidents.md` by date.

Built from the best public work on this

Sources for supplier-scorecard

Everything below was opened and read on 13 September 2026. Nothing is cited that could not be loaded.

1. Wikipedia, "DIFOT"

https://en.wikipedia.org/wiki/DIFOT, last edited 20 July 2026.

The encyclopaedia article on Delivery In Full, On Time, the metric more commonly written as OTIF. It supplies the arithmetic the skill uses in step 3, stated on the page as "OTIF (%) = number of OTIF deliveries ÷ total number of deliveries × 100", and it lists the three things a business needs before the sum means anything: a specified delivery date on the order, a documented actual delivery date in the records, and a recorded reason for each failure. That list is why step 1 refuses to score a delivery with incomplete paperwork rather than guessing at it, and why step 2 insists every cell carries a document number. The article also notes the calculation may be run "by orders or order lines rather than deliveries", which is the choice step 4 makes explicit: a restaurant feels a missing line, so the skill reports line fill rate and whole-order perfection side by side instead of picking one. Most valuable is the limitation the article records from research by Janet Godsell, that suppliers frequently meet OTIF targets when measured against the delivery dates they themselves promised and miss when measured against the dates the customer requested. The skill turns that into a hard requirement to record both dates on every row. Where the skill goes further than the source: the article is written for supply chain professionals with an ERP behind them, and simply assumes the delivery date is in a system. For a 40-cover restaurant it very often is not, so the skill spends its first step on reconstructing the promised date from the order confirmation or the owner's own written note of a phone order, and treats an unreconstructable delivery as an excluded row with its own count, rather than as a scoring problem.

2. GoodSource, "Vendor Performance Evaluation Metrics for Wholesale Food Distribution Partnerships"

https://goodsource.com/trends-and-insights/vendor-performance-evaluation-metrics-for-wholesale-food-distribution-partnerships/, 22 March 2026.

A foodservice procurement article aimed at operators buying from wholesale food distributors, which is exactly the relationship this skill measures. It is the only source here written for the specific failure modes of food delivery rather than for factory components. Three things came from it. First, the substitution example in step 5: the article describes a distributor replacing a 10-pound case of chicken breasts with two 5-pound cases at the same per-pound price, and argues the standards for what counts as an acceptable swap must be set before it happens. Second, the price consistency band in step 6, where the article treats a 3 to 5% variance as the range within which pricing can be called consistent. Third, the cadence discipline in step 9, where it recommends monthly scorecards go out within 48 hours of the period closing on the basis that delayed feedback loses its force. The skill disagrees with the source in two places, and both disagreements matter. The article proposes a weighting of roughly 30% delivery, 25% quality, 20% cost and 25% for communication, problem resolution and partnership cooperation. The skill cuts the soft quarter to 10% for credits and admin and shifts the weight onto order accuracy and substitutions, because "partnership cooperation" cannot be evidenced from an invoice file and a category scored on feel is the category an account manager will argue you out of. The skill also drops the article's American cold chain figures, which are given in Fahrenheit as 32 to 40°F chilled and 0°F frozen, and uses the UK legal limit instead: an English kitchen that rejects a load on an American threshold has neither a legal basis for the rejection nor a defensible record for its environmental health officer. Finally, the 48 hour turnaround is relaxed to five working days, because it is written for a procurement team and this pack is for an owner who is also on the pass on Saturday night.

3. EvaluationsHub, "Supplier Scorecard Best Practices: KPIs, Weighting, Cadence"

https://evaluationshub.com/supplier-scorecard-best-practices-kpis-weighting-cadence/, 4 March 2026.

A practitioner guide to designing scorecards rather than to any one industry. Its most useful contribution is structural: it argues for a master KPI catalogue that defines every measure "in simple, auditable terms" before any scoring begins, for normalising everything to a common scale such as 0 to 100, and for explicit rules covering missing data and outliers. Step 2 of the skill is that idea applied to a shoebox of invoices, which is why the column list is fixed and closed rather than left to the owner. The guardrail line quoted in step 9, that weights should be set "so one KPI cannot dominate the total score", is the reason the skill caps every category at its own weight and refuses to let a keen price buy back a bad delivery record. Its cadence table, which puts strategic partners on monthly scorecards with quarterly business reviews and transactional suppliers on quarterly scorecards with annual reviews, is behind the choice of a 13 week window rather than a monthly one. Where the skill parts company with the source: the guide assumes the data arrives from an ERP and a CRM, naming those systems as the connection points, and its dispute handling is a "structured feedback loop" between two procurement functions. Neither exists in a pub. The skill therefore treats the paper documents themselves as the system of record and makes the document reference the dispute mechanism, on the basis that when an account manager challenges a figure the owner needs to be able to put the credit note on the table in ten seconds rather than promise to look into it.

4. GOV.UK (Food Standards Agency guidance), "Food hygiene for businesses: Chilling and freezing"

https://www.gov.uk/food-hygiene-businesses/chilling-and-freezing, no publication date shown on the page, read 13 September 2026.

The official UK guidance page for food businesses on chilled and frozen storage. It states plainly that "Chilled food must be kept at 8°C or below. This is a legal requirement in England, Wales and Northern Ireland", and follows it with the working recommendation to "Set your fridge to 5°C or below to make sure food stays cold enough, even if the temperature changes." Those two figures are the whole of step 8: the 8°C line is the point at which a delivery fails and the rejection has to be recorded, and the 5°C target is what the owner should actually be asking a supplier to deliver at so there is margin for the walk from the lorry to the walk-in. This source is the reason the skill treats temperature as a separate scored category rather than folding it into a general quality score. A temperature failure is not a service disappointment to be averaged; it is a breach the owner must be able to show they caught and acted on. Note the limit of the citation: the page is written about the business's own storage and freezing practice, not about grading a supplier, and it does not set out goods-in procedure. The skill is extending its figures to the moment of delivery, which is a reasonable reading given the same legal limit applies to the food on arrival, but it is an extension and that is why the final rule sends the output to the owner's environmental health officer rather than treating the page as authority for a rejection policy. Scotland is deliberately not covered by the quoted wording and the skill does not claim it is.

5. asgard-ai-platform, "mfg-supplier-scorecard/SKILL.md"

https://github.com/asgard-ai-platform/skills/blob/main/mfg-supplier-scorecard/SKILL.md, 228 stars read from api.github.com, no publication date shown on the file, read 13 September 2026.

A published agent skill for supplier scoring in manufacturing, part of a repository of 301 such skills. It is the closest public artefact to the job this skill does and the structural backbone came from it: four weighted dimensions (Quality 30 to 40%, Cost 20 to 30%, Delivery 20 to 30%, Service 10 to 20%), a defined scale rather than a vibe, a weighted total, and grade bands A above 4.0, B from 3.0 to 4.0, C from 2.0 to 3.0 and D below 2.0, each attached to an action rather than just a label. The C band's "improvement plan required within 90 days" is the response step 9 recommends at a renewal, because for a single-site restaurant the realistic outcome of a bad scorecard is a harder conversation and a review date, not switching wholesaler mid-season. Two clear disagreements. First, the file scores each KPI on a 1 to 5 scale with descriptions such as "top 10% of suppliers", which requires a supplier base to rank against; a restaurant usually has one produce supplier and cannot be in a top decile of one, so the skill scores every category out of 100 against absolute targets taken from the owner's own paperwork instead. Second, its quality dimension is built on defect rates in parts per million and incoming inspection pass rates, which do not survive translation to a crate of tomatoes. The skill replaces that dimension entirely with order accuracy, substitutions and delivery temperature, which are the things that actually go wrong in a kitchen and, unlike PPM, can be evidenced from documents the owner already has in a drawer.

Best public prompt we found for this job

The single best public artefact is the `mfg-supplier-scorecard` SKILL.md in the `asgard-ai-platform/skills` repository (228 stars, read from api.github.com), which opens its framework with an "IRON LAW" block:

The cheapest supplier who delivers defective parts late with no support
is the most expensive supplier.

That line is worth copying because it fixes the one thing an owner gets wrong at renewal: the price on the quote is the only number that is easy to see, so it wins the argument by default, and every cost of a bad supplier (the missing item at 7pm, the swapped product that broke a menu description, the credit chased for six weeks) sits outside it and never gets counted. The whole point of the scorecard is to put those costs into pounds so they can stand next to the price.

Want this running in your business?

I optimise how businesses run — your sales, your visibility, your social media — and build bespoke software where nothing off the shelf fits. The first conversation is free. Work starts from £150 a day.