01

Fictional project · real numbers · plays in about ninety seconds

A project, played back

Seven delivery calls on one project, played back with what each one cost. The project is an automation rollout in a glass plant: 26 weeks, €400,000 and three driverless vehicles taking pallets across an internal street, in a plant that cannot stop while the work goes in. The plant and the people are made up; the numbers come from a real project charter, and I ran a production line in a plant like this one. It ends on the number that decided the business case: production lost while the project was being delivered.

Step 1 of 13 · week 0 of 26
Week 00/26
Capex spent €0k / €400k
Production lost €0k
Payback 2.5 yrs
Crew buy-in 40%
Safety case 55%

Payback is capex plus production lost during delivery, divided by the €160,000 a year the project saves. Production lost is the number to watch.

Week 0 · Kickoff

Five lines, four warehouses and the street between them.

Every pallet made in this plant is driven across an internal street by a forklift. Twenty-six weeks and €400,000 to stop doing that on one route, without stopping anything else.

Week 2 · Call 1

The board wants all five lines

The payback case was good enough to raise an obvious question. Three vehicles cover one route. Five lines needs thirteen and a second corridor.

The reasoning is below, under the timeline.

Week 6 · Call 2

Safety won't sign a corridor with one layer of control

The route runs the length of the cold end and out through the warehouse door, the same door the shrink wrap operators use to reach the crib room.

The reasoning is below, under the timeline.

Week 9 · Call 3

A dead zone exactly where they pick up

Fourteen metres between the strapper and the shrink tunnel with no signal when the racking is full. Metal, steam, and pallets stacked three high.

The reasoning is below, under the timeline.

Week 11 · Install

Route marked, chargers in, vehicles on site

Floor markings, barriers at the pinch points, two chargers on the warehouse side. Installed around a plant that never stopped running.

Week 12 · Call 4

"What happens to my six?"

Someone printed the business case and pinned it to the notice board. It says reduce forklift traffic by seventy percent, and the drivers can do that arithmetic.

The reasoning is below, under the timeline.

Week 16 · Call 5

02:40, the same pallet, refused four times

Out of profile by thirty millimetres. Every rejected pallet came off the shrink lane, none off the strapper, and night shift's answer was to send a forklift.

The reasoning is below, under the timeline.

Week 18 · Integration

Fleet manager talking to the warehouse system

Missions dispatched from the palletiser signal, pallets confirmed by RFID on arrival, locations written back to the WMS.

Week 20 · Call 6

A €38,000 change order, four weeks out

The warehouse system expects a confirmation per pallet. The fleet manager posts one per mission. The signed specification says per pallet.

The reasoning is below, under the timeline.

Week 23 · Testing

Eleven hundred test moves, and a rehearsed way back

Safety case signed. Drivers trained as marshals. Manual fallback rehearsed twice: eighteen minutes to put forklifts back on the route.

Week 24 · Call 7

Vehicle 3 keeps losing itself in the same spot

Two localisation faults in a week, both within six metres of each other, both self-recovered. The go-live window is 48 hours and the next one is in March.

The reasoning is below, under the timeline.

Week 25 · Go live

Two vehicles live. The third follows.

Map re-taught near the tunnel, vehicles 1 and 2 running the route, marshals getting a fortnight at a pace they can handle.

Week 26 · +90 days

The forklifts stopped crossing the street.

All three vehicles live. Fifty-four percent of driver hours moved to work that needed a person. And the shrink tunnel has been running at temperature since week sixteen.

Same project, delivered the other way: 10.7 years ↓

Schematic of a container glass plant: two furnaces feeding five forming lines, glass running through the cold end, wrapping lanes and palletisers, then crossing an internal street into the warehouses. WAREHOUSE A WAREHOUSE B WAREHOUSE C WAREHOUSE D INTERNAL STREET CHARGERS PALLETISERS WRAPPING LANES INSPECTION LINE 1 FORMING LINE 2 FORMING LINE 3 FORMING LINE 4 FORMING LINE 5 FORMING FURNACE 1 FURNACE 2 RAW MATERIALS FLEET SOFTWARE · WAREHOUSE SYSTEM PROJECT OFFICE

Glass runs right to left: furnaces, forming lines, inspection, wrapping, palletisers, then across the internal street to the warehouses. The vehicles' route is the street.

  • Running
  • Under pressure
  • A problem
  • Delivered
Call 1, week 2: what I chose and why
Email · Marta Ferrer, Operations Director

David,

The board liked the payback. 2.5 years on €400k is the best capex case we've put up this year, so they've asked the obvious question: if it works for one line, why are we only doing one line?

I'd like the charter redrawn for all five. Same budget if possible.

Marta

What I chose

Deliver route one as chartered, and hand Marta a costed three-phase expansion with the numbers behind it.

She gets a board answer (yes, here is what all five costs and in what order) without you committing €400k to €1.4m of work. Phase 2 gets a slot in next year's capex round on your figures. The pedestrian crossing on the Warehouse B route becomes a named phase 3 problem instead of an unnamed phase 1 one.

Marta's real question was what to tell the board, and that has a far cheaper answer than five lines. A costed roadmap is two days of work and it converts a scope fight into a planning conversation.

What the AI gave me: What three vehicles can carry, against what five lines produce.
In

Charter and business case, 12 months of pallet despatch data by line, cold end layout, shuttle car cycle times, the vendor's throughput specification.

Out
RoutePallets/h peakAGV round tripVehicles neededCorridor share
Line 3 → Warehouse A146.2 min2.1low
Lines 1 and 2 (amber) → Warehouse A267.8 min4.9high
Lines 4 and 5 → Warehouse B2111.4 min5.6crosses pedestrian route
All five61n/a12.6n/a

Three vehicles cover one route at peak with 30% headroom. Five lines needs thirteen, two more chargers and a second corridor: roughly €1.4m, against a €400k charter.

Checked

The 6.2 minute round trip is the vendor's figure and it assumes a clear run. I timed the corridor at shift change with a stopwatch: 8.9 minutes with pedestrian stops. Re-ran the model on the real number and the headroom on one route drops from 30% to about 8%. That changes my answer on vehicle count even for the single route. Better to find that in week 2 than in week 24.

Time

35 minutes, most of it standing in the corridor with a stopwatch. The model took two.

The options I didn't take, and why
Redraw the charter for all five lines within the same budget.

Saying yes to a scope you have already modelled as undeliverable is a decision to fail later, in public, on your own signature.

Say no. The charter is the charter: one route, three vehicles.

Right answer, wrong delivery. The sponsor asked for five lines because the board asked her. If I hand back a flat no, she carries it alone, and I need her in week 20.

Take two routes, line 3 plus the amber lines, and ask for €250k more.

Reopening an approved budget is expensive in a way that doesn't show up on the budget. And picking the highest-traffic corridor for the first deployment is choosing to learn on the hardest case.

Call 2, week 6: what I chose and why
In person · Aiko Tan, Safety Engineer

"Your route runs the length of the cold end and out through the warehouse door. That's the same doorway the shrink-wrap operators use to get to the crib room, and the same aisle Rob's fitters walk down with toolboxes.

ISO 3691-4 doesn't say you can't mix people and vehicles. It says you have to show me you've assessed it. Right now the design has scanners on the vehicles and nothing on the corridor.

I'm not signing a safety case that puts the entire control on the thing that's moving."

What I chose

Barrier at the three pinch points, marked walkway elsewhere, and a door interlock on the crib room.

€22k and two weeks. Aiko co-writes the risk assessment rather than reviewing yours, and signs it. The interlock stops a vehicle whenever the crib room door opens, which is where 61% of the exposure actually is.

The survey is what makes this the right answer instead of a guess. Two spikes a shift at one door means you put the hard control on the door. And getting Aiko to write the assessment turns the person with the veto into the person with authorship.

What the AI gave me: Four segregation options against cost, floor space and what each one controls.
In

ISO 3691-4, vendor safety spec, cold end layout, pedestrian route survey over four shifts, incident log 2019 to present.

Out
OptionCostSpaceControlsResidual
Vehicle scanners onlyin basenoneVehicle stops for a detected personHigh: relies on detection alone
Physical barrier full length€68k1.4 m, and the corridor won't take itSeparationLow
Barrier at pinch points + marked walkway€22k0.9 m at 3 pointsSeparation where it's tight, marking elsewhereLow-medium
Zone speed limiting + door interlock€9knoneVehicle slows in shared zones, stops at the crib room doorMedium

Pedestrian survey, four shifts: 214 corridor transits. 61% are the crib room door at break. 22% are fitters between the wrapping lanes and the workshop. The rest is spread.

Checked

The model treated the 214 transits as evenly distributed and recommended zone speed limiting on that basis. They are two spikes, twenty minutes long, twice a shift. That makes a door interlock and a hard barrier at two pinch points far more effective than slowing the vehicle for the whole shift, and cheaper than the full barrier. The averaged number would have given me the wrong control.

Time

The options table, 15 minutes. The transit survey was four shifts of someone standing there with a counter, and it is the only reason the answer is right.

The options I didn't take, and why
Vehicle scanners only. The vehicles are certified. Argue the case.

The vehicle being certified is not the same as the installation being safe. She is telling me my design has one layer of protection and she needs two.

Full physical barrier down the corridor.

The safest option is not automatically the right one when it cannot physically fit. I'd rather bring Aiko three options that exist than one that doesn't.

Zone speed limiting across the shared length, no barriers.

This is what the averaged data told me to do, and it is the option I'd have taken if I hadn't gone and counted. Controlling for the average when the risk is a spike costs you throughput permanently and still leaves the spike.

Call 3, week 9: what I chose and why
Teams · Priya Raman, IT Integration

Wi-Fi survey back. Corridor and warehouse are fine, strong signal the whole run.

The problem is the wrapping lanes. Between the strapper and the shrink tunnel we've got a 14 metre stretch reading -78 dBm and dropping out entirely when the racking either side is full. Metal, steam off the tunnel, and pallets stacked three high acting like a wall.

The AGVs need continuous coverage to hold a mission. In that stretch they'll go dead-reckoning and then stop.

That stretch is where they pick up.

What I chose

Two ceiling-mounted access points, cable run over the shrink tunnel.

€6k and three weeks, and it holds when the racking is full because the coverage comes from above rather than through the stack. Working over a live shrink tunnel needs a permit and a hot work assessment.

Fix the physics. Coverage from above is the only version that survives full racking, and full racking is the normal state of that area.

What the AI gave me: Three ways to fix a dead zone, and what each does to the commissioning date.
In

Wi-Fi heat map (4 passes, loaded and empty racking), AGV vendor comms spec, plant network topology, access point inventory.

Out
FixCostLead timeCoverage afterNote
Two extra access points, ceiling mount€6k3 weeks−58 dBmNeeds a cable run over the shrink tunnel
Directional AP each end of the stretch€4k2 weeks−64 dBm loadedDegrades when racking is full
Private 5G for the cold end€90k16 weeks−52 dBmSolves it permanently, misses go-live
On-vehicle mission bufferingvendor says supportedn/an/aVehicle completes mission through a dropout
Checked

The fourth row is the one that mattered and the model took the vendor datasheet at face value. I asked Lucas directly: mission buffering holds a mission through a dropout of up to 8 seconds. Our dead stretch takes 40 seconds to traverse at survey speed. So the feature is real, and it does not solve our problem. Datasheet said supported; the number underneath it said no.

Time

20 minutes. The question that saved the project was 'supported for how long', and it isn't in the table.

The options I didn't take, and why
Trust the vendor's mission buffering. No network work.

'Supported' is a marketing word until you make it a number. Eight seconds of buffering against a forty second dead stretch is a coincidence waiting to fail.

Directional APs at each end. Cheaper, faster, no work over the tunnel.

The trap here is that it will pass its acceptance test. You'll sign it off in a quiet week and inherit an intermittent fault that takes three months to diagnose because it only happens when you're busy.

Private 5G for the whole cold end. Solve it properly and permanently.

This is the answer if we were designing the plant. We're not, we're inside a project with a window, and a solution that arrives after the window is not a solution.

Call 4, week 12: what I chose and why
In person · Dan Toohey, Shift Team Leader

"They've read the business case. It's on the notice board. Someone printed it.

'Reduce manual forklift traffic by seventy percent.' Two or three drivers a shift, three shifts a day. They can do that arithmetic and so can you.

I'm not here to stop your project. I'm here because they're going to hear something in the next fortnight and I'd rather it came from you, with a number in it, than from the rumour that's already going round.

What happens to my six?"

What I chose

Give him the real number in writing: 54% of hours, not 70%, and what stays.

It is a better number than the one on the notice board and it is checkable, which is why it lands. Nearly two hundred hours a week of dock loading, cullet moves and mould shop runs still need a driver, and AGV marshalling is new work for someone who already knows where pallets jam. Dan takes it to his crew himself.

The honest number was better than the rumour, which is usually the case and almost never the reason people withhold it. And the correction matters beyond this conversation: a business case that says 70% when it means corridor traffic will keep generating this problem until someone fixes the wording.

What the AI gave me: Where six drivers' hours go, and what the AGVs do and don't take.
In

Forklift telemetry 12 months, shift rosters, task observation over 6 shifts, warehouse vacancy list, business case assumptions.

Out
TaskDriver hours/wkAGV takes it?
Wrapping lane → Warehouse A run227Yes. This is the 70%
Container and cullet bin moves55No. Not on the route, not palletised
Loading outbound trailers71No. Needs a driver at the dock
Mould shop and spares moves38No. Ad hoc, no fixed route
Recovering fallen and damaged pallets29No. And this goes up in year one

Around 420 driver hours a week across three shifts. The AGVs take 227 of them: 54%, not 70%. The remaining 193 hours still need drivers, plus new AGV marshalling and fault recovery at an estimated 12 hours a week.

Checked

The business case's 70% is a reduction in corridor traffic, not in driver hours, and somewhere between the model and the notice board those became the same sentence. They are not. I checked the telemetry myself rather than the case: 54% of hours, and the residual tasks are the ones that need a person. That distinction is the entire answer to Dan's question, and the business case wording is what created the problem.

Time

40 minutes. The number was never the hard part. Noticing that two different percentages had been written as one was.

The options I didn't take, and why
Tell him it's an HR matter and route it through the process.

A question asked in front of six people gets answered in front of six people or you have answered it anyway, badly.

Promise no redundancies.

Never promise something you can't deliver to buy peace in a difficult meeting. The relief lasts a fortnight and the damage lasts the rest of your time there.

Bring the drivers into commissioning as the AGV marshals.

I'd do this as well, not instead. It's the best thing you can do for the project and for them. But it is a job offer to some of them, not an answer to all of them, and Dan asked a numbers question.

Call 5, week 16: what I chose and why
Alarm · Fleet manager, fault log, 02:40

Vehicle 2 has refused the same pick-up four times tonight. Vehicle 1 took it on the fifth attempt.

The fault is PALLET_GEOMETRY_REJECT. The forks won't engage because the load is out of profile: the wrap has pulled the top courses in and the bottom is proud by about 30 mm on one side.

Every one of the four came off the shrink lane. None off the strapper.

Night shift's answer was to send a forklift. Which works, and quietly makes your automation optional.

What I chose

Fix the shrink tunnel. Get the exit temperature holding overnight.

A burner control fault and a door seal, found in a day once someone was looking. Reject rate on the shrink lane drops from 8.7% to 1.4%. Transit damage claims fall too, which nobody predicted and nobody had been measuring.

Fix the thing that is broken. The AGVs exposed a fault the plant had absorbed by hand for a decade. The payback from the tunnel fix showed up in a completely different budget line, which is the honest reason automation projects are hard to appraise up front.

What the AI gave me: Six weeks of pallet rejects, sorted by which wrapping lane they came off.
In

Fleet manager fault log (6 weeks), lane despatch records, shrink tunnel temperature trend, pallet profile spec from the vendor.

Out
LanePalletsRejectsRateDominant fault
Strapper2,410190.8%Skewed base board
Stretch wrap1,884412.2%Film tension, top courses drawn in
Shrink tunnel1,102968.7%Out of profile, one side proud
Strap + wrap combined64071.1%n/a

71% of shrink lane rejects fall between 01:00 and 05:00, and they correlate with tunnel exit temperature running 8 to 11 °C below setpoint on night shift.

Checked

The obvious read is 'the AGVs are too fussy, widen the tolerance'. The night-shift clustering says otherwise: this is a shrink tunnel running cold at night and not shrinking the film properly. The AGV is the first thing in eleven years that has been unable to ignore it. Widening the tolerance would have hidden a real process fault that was already costing us in transit damage.

Time

18 minutes to find the correlation. Rob knew the tunnel ran cold at night. Nobody had ever had a reason to write it down.

The options I didn't take, and why
Widen the AGV pallet tolerance so it accepts what the lanes produce.

Turning off an alarm is not the same as fixing what set it off. The vehicle found a real defect that the plant had been absorbing manually for a decade. That is the automation earning its money.

Route shrink-lane pallets to a manual forklift, automate the other three.

This is a scope cut dressed as a routing decision. It also parks a driver next to the defect so nobody ever has to deal with it.

Escalate to the vendor. Their forks should handle a real pallet.

Check whether you're inside your own spec before you accuse a supplier of missing theirs. I've been on the wrong end of that conversation and it costs you more than the week.

Call 6, week 20: what I chose and why
Email · Lucas Meyer, vendor delivery manager

Hi David,

During integration testing we've found the WMS expects a location confirmation per pallet, but our fleet manager posts a mission-complete per mission. On a double-handling move that's two pallets against one message.

Aligning this needs development at our end. We've scoped it at 24 days, €38,000, and we'd need sign-off this week to protect the week 24 date.

Lucas

What I chose

Hold the §4.2 position, and offer to take the per-pallet split on our side of the interface for a schedule guarantee.

Priya can split mission-complete into per-pallet confirmations in our middleware in about three days, because the RFID reads are already in the message. Lucas keeps his margin, we keep the €38,000, and in exchange he puts the week 24 date and two on-site engineers in writing.

The contract question and the delivery question have different best answers, and you don't have to win both. I have the stronger legal position and the weaker practical one. I need his engineers more than I need €38,000. So I trade the position I'd probably win for something I actually need, and he gets to report a solved problem rather than a lost argument.

What the AI gave me: What the interface specification says, and who owns the gap.
In

Signed interface specification, vendor SOW, WMS message schema, integration test logs, two comparable deployments' rate cards.

Out
SourceSays
Interface spec §4.2, signed by both"Mission completion shall post a confirmation per pallet including pallet RFID, destination and timestamp."
Vendor SOW, deliverable (f)"Integration with customer WMS in accordance with the agreed interface specification."
Fleet manager, as deliveredPosts per mission. Does not meet §4.2.
Comparable deployment rate€1,150/day. CO-scoped at €1,583/day, 38% above.

Reading it plainly: the specification both parties signed says per pallet. The delivered system does not do that. This is a defect against an agreed specification, not a change in scope.

Checked

I read §4.2 myself rather than trusting the summary, because this is the kind of finding that is satisfying and occasionally wrong. It says what the model said it says. But I also checked whether we ever asked for double-handling moves. We did, in week 11, and it is minuted. So the requirement is ours, it is in the spec, and it predates the CO. That sequence is what makes this a defect rather than an argument.

Time

25 minutes. Twenty of them were finding the week 11 minute that proves the requirement wasn't invented after the fact.

The options I didn't take, and why
Pay the €38,000 from contingency to protect the date.

Sometimes buying the date is right. Not when the thing you are buying is already in the contract, and not at week 20 when you still have the riskiest fortnight in front of you.

Reject it as a defect under §4.2 and require remediation at their cost.

The contractual position is correct and I'd hold it. What I wouldn't do is hold it in a way that leaves Lucas nothing to take back to his own management.

Accept mission-level confirmation and change the WMS instead.

Never solve a supplier's integration problem by weakening your own system of record. The cost lands on people who were never in the project.

Call 7, week 24: what I chose and why
In person · Go-live decision, 15:00 the day before

Everything is staged. Safety case signed, drivers trained as marshals, chargers commissioned, WMS interface passing.

The go-live window is the low-season shutdown: 48 hours, and it is the only window until March.

Two things are unresolved. Vehicle 3 has thrown two unexplained localisation faults in the last week. Both recovered themselves in under a minute, both in the same spot near the shrink tunnel. And the marshals have had four days of hands-on, not the ten in the plan, because the vendor engineer arrived late.

Marta wants all three vehicles live tomorrow. It is on a board slide.

What I chose

Re-teach the map near the tunnel, then go live on vehicles 1 and 2. Vehicle 3 follows in two weeks.

Four hours of re-teaching clears vehicle 3's fault. Two vehicles cover the route at 82% of peak, which the shuttle cars absorb. The marshals get a fortnight of real operation at a manageable pace, and vehicle 3 joins with the map fixed and its own week of watching.

Go-live is a sequence, and the sequence is your last chance to buy risk back. Two vehicles with confident marshals beats three with a four-day crash course, and the difference shows up in whether anyone reaches for a forklift in week one. Marta loses a line on a slide. She would lose considerably more from a fortnight of firefighting.

What the AI gave me: A readiness scorecard and a pre-mortem: it is ninety days from now and this failed.
In

Commissioning test logs, safety case sign-off, marshal training records, localisation fault dumps, charger utilisation, six weeks of dry runs.

Out
GateStatusEvidence
Safety case, ISO 3691-4PassSigned wk 22, interlock and barriers verified
Vehicles 1 & 2 localisationPass0 faults, 14 days
Vehicle 3 localisationWatch2 faults, same location, self-recovered
WMS interfacePassPer-pallet confirmation, 1,100 test moves
Marshal trainingWatch4 days of 10. Vendor arrived late
ChargingPassTwo chargers sustain three vehicles at peak
Manual fallbackPassForklift recovery rehearsed twice, 18 min

Pre-mortem: it is ninety days on and this failed. Why?

  1. Vehicle 3's fault was the new ceiling access points changing the feature map near the tunnel, and it degrades as racking fills. Counter: re-teach the map in that stretch before go-live.
  2. A marshal couldn't recover a stopped vehicle at 3am, sent a forklift, and the workaround became the process. Counter: don't go live wider than the trained marshals can cover.
  3. The shrink tunnel drifts cold again and rejects climb with nobody watching the lane split. Counter: put the per-lane reject rate on the shift board.
  4. Nobody owns the fleet manager after the vendor leaves. Counter: name it before go-live, not after.
Checked

Failure 1 is the one that decides tomorrow. Both vehicle 3 faults are within 6 metres of each other and both are after week 9, which is when we put the new access points in over the shrink tunnel. That is a stale feature map in a stretch we physically changed. Re-teaching that section is four hours, and it explains the symptom completely. I would not have connected those two facts without the fault dump coordinates.

Time

30 minutes. The pre-mortem was going to happen anyway; the coordinates are what turned a mystery into a four-hour job.

The options I didn't take, and why
All three vehicles live tomorrow, as briefed to the board.

This is the failure I have watched happen, and the technology was fine. The moment you take the manual route away before the automated one is proven, a one-hour fault becomes lost production on every line behind it. Labour savings of a hundred and sixty thousand a year do not survive that, and the loss lands in a completely different budget line from your capex, so nobody sees it coming until the plant number comes out.

Postpone. Take the March window with everything resolved.

Postponing to fix a four-hour map re-teach and six days of training is an inability to sequence. If the safety case were open I'd take March without blinking.

Go live on all three, keep two forklifts running the route in parallel for a month.

Parallel running is genuine risk control in a data cutover, where the old system costs nothing to leave switched on. Here the old system is people, and running both tells them which one you believe in.

02
The same project, delivered the other way

Same vehicles, same vendor, same plant. Delivered this way the payback is 2.6 years. Delivered the other way it is 10.7 years, and the capex barely moves.

Delivered this way Delivered the other way
Capex €313k €347k
Production lost to delivery €98k €1365k
Payback 2.6 years 10.7 years
Crew buy-in at go-live 94% 2%
Safety case 100% 25%

Every choice in the second column is defensible in the room. Take the scope the board asked for. Put the hazard behind the biggest barrier available. Trust the vendor's datasheet. Send the staffing question to HR. Widen a tolerance that is rejecting pallets. Go live on everything at once, on the date you promised. What moves is production lost while installing and cutting over into an operation that cannot stop. The business case was built on labour savings of €160,000 a year, and the production lost in the second column is worth more than 8 years of them.

The lesson

The delivery method is the business case. That is the whole lesson, and it is the one I did not properly understand until I watched a project like this go wrong.

A payback case built on labour savings is a thin thing. €160,000 a year is about €3,000 a week. So a single shift lost across five lines, caused by your own installation or your own cutover, costs more than a month of the benefit you are there to deliver. A fortnight of disruption costs more than three years of it. And it does not appear anywhere in your capex tracking, because lost production lands in the plant's number, not the project's.

I have seen a rollout of this shape take its payback from months to years on exactly that: production lost while it was being installed and switched on. No overspend, no technology failure. The capex came in fine. The project still destroyed its own case.

In a plant that cannot stop, three things become the project:

Sequence so there is always an unaffected line. The board's instinct is to do everything at once because it looks efficient. Installing across the whole cold end simultaneously means every line takes disruption together and nothing is left to carry the plant. Phasing is what keeps the loss bounded when something goes wrong, and something will.

Never remove the manual route until the automated one has proven itself. This is the single decision that separates the two columns above. Once the corridor is reconfigured for automated traffic and the forklifts cannot easily take back over, an ordinary one-hour fault turns into lost production on every line behind it. Keeping the old route rigged and drivable costs almost nothing and converts every unknown from a loss into an inconvenience.

Hold contingency for production impact as well as capex. Every project holds a cash contingency. Almost none hold a production contingency, and on a site like this that is the exposure that is an order of magnitude larger. If you cannot say what a day of lost output on your busiest line costs, you cannot price any of your own delivery decisions.

The other four calls in the run matter and I would defend each of them: sequencing the scope instead of cutting it, putting the hard safety control where a survey said the people were, giving the drivers the real 54% instead of the notice board's 70%, and fixing a shrink tunnel the automation had just exposed. None of them would have saved this project from a big bang go-live. That one decision was worth more than the other six put together.

The criticism you would still take: one route live out of five. The answer is that the other four are costed and sequenced in the capex round, and this one paid back. That is a better position than five routes and a plant number nobody wants to discuss.

Where the AI was

Each of the seven calls above has an analysis behind it, drafted with AI in under an hour.

None of them made a decision. Two were wrong in ways that mattered: the corridor model averaged 214 pedestrian transits that were really two spikes a shift, and the network options table repeated a vendor's word "supported" without the eight-second number underneath it that made the feature useless to us. Both were caught by going and looking: four shifts with a counter, and one direct question to the vendor.

That is the job now. The analysis is cheap, so the value moved to framing the question, knowing which inputs are missing, and owning the call.

Nothing here identifies a site. The layout is a composite of container glass plants, and the plant, the people and the vendor are fictional. The constraints and the failure modes are real.