CASE STUDY · MACHINE LEARNING, BUILT AND DELIVERED
One expert's knowledge, now a system the plant can use
A roll manufacturer could not set a heat-treatment recipe without its founder. We built the machine learning system that moved that knowledge into software, and raised recipe accuracy from 44% to 92.7%.
Sector | Manufacture and reconditioning of forged and hard-surfaced rolls for steel, cement and forging |
What we built | Heat Treatment Assistant — an ML-based web application |
Project type | New build |
Timeline | Mid-June 2026 to end of July 2026. Live 7 July 2026. |
Still running | Yes — 3 months, last production use 25 September 2026 |
Stack | Python / FastAPI, scikit-learn, React 18, PostgreSQL, Docker on the client's own server |
The problem
Every heat-treatment recipe at the plant depended on one person. The founder was the only one who could verify the parameters, so every job queued behind his availability. The knowledge was real and hard-won, but it existed in one head and nowhere else.
CST is a manufacturer, not a software company, so building it internally was never an option. The risk was simple and getting worse: a single point of contact on a step that every job has to pass through.
What we built
The Heat Treatment Assistant takes the details of a roller job — steel grade, required hardness, diameter, roll type and furnace — and returns the full recipe: soaking temperature, soaking hours, heating rate and cooling rate. Any authorised engineer can use it, from anywhere, behind two-factor login.
Three decisions that shaped it
· The model chooses from settings the plant has actually run, rather than predicting a raw number. A furnace runs at 550 or 580, never 551.3. It can never suggest a setting the plant has never used.
· One model per output, each using only the inputs that measurably improve it. Inputs that tested as noise — quantity, job weight — were removed rather than kept for appearance.
· Accuracy is measured with grouped validation, so a repeated job can never sit in both training and test. This is why the improvement figure is real rather than flattering.
What we considered and rejected
· Regression, predicting raw numbers. Rejected: the answer was never exactly right and it could suggest settings outside the plant's real practice.
· Deep learning. Tested and rejected: on a few hundred job records neural networks performed no better than tree ensembles, and are heavier to deploy and maintain on a plant server.
· Direct lookup of past jobs. Rejected: it scores perfectly on jobs already seen and fails on anything new.
What was genuinely hard
The data contradicted itself. The same job appeared with different settings, and both had passed. We traced the cause — recipes are set per furnace load, and standards had drifted over the years without being recorded — then resolved every job to the standard the plant uses today, and exported the conflicting ones to a separate sheet for the plant to confirm.
The records were messy. The same grade spelled many ways, mixed units, date typos. Of 649 rows, 455 were trustworthy. Every removal has a logged reason, so the training set can be audited rather than taken on trust.
Proving the accuracy honestly was its own problem. Repeated jobs make flattering numbers easy. We measured the old engine the same way for a fair before and after, and verified the deployed app against live predictions before release.
What the client can do now that they could not before
· Set a heat-treatment recipe without waiting for one person. The knowledge that lived in a single expert's head is now available to any authorised engineer in seconds.
· See the evidence behind every recipe. Each recommendation carries its confidence, the nearest alternative, and a check against what that grade actually achieved in past production.
· Audit and improve over time. Every prediction is stored with its outcome, and new job records are imported monthly to retrain, so any answer can be traced back to real jobs.
· Test changes safely. When their data reviewer asked what would happen if job weight were added, we shipped it as a side-by-side version in the live app and measured it. The question was answered with evidence.
How it was tested and delivered
Automated tests | 73 backend, 12 frontend, plus Playwright end-to-end against a real login |
Pre-release check | A server simulation run from exactly the files being shipped, before every push |
Final verification | Live predictions checked against the plant's own completed job records on UAT |
Deployment | Docker on the client's own server. No cloud, no external services. Access via Cloudflare Tunnel with two-factor login. |
Client sign-off | The client's own data reviewer challenged the result. We measured his alternative live in the app. The client signed off. |
Automated tests
73 backend, 12 frontend, plus Playwright end-to-end against a real login
Pre-release check
A server simulation run from exactly the files being shipped, before every push
Final verification
Live predictions checked against the plant's own completed job records on UAT
Deployment
Docker on the client's own server. No cloud, no external services. Access via Cloudflare Tunnel with two-factor login.
Client sign-off
The client's own data reviewer challenged the result. We measured his alternative live in the app. The client signed off.
How our engineers used AI on this project
We publish this because buyers now ask, and most vendors cannot answer it.
Tool used
Claude Code, throughout development
In the product
None. No AI service runs inside the delivered application and no plant data passes through one. The AI was in how we built it, not in what we shipped.
Client policy
The client had no AI policy. AI tools were permitted for development work.
Review
Every change passed the full test suites and live verification, and a developer approved every commit and push. Nothing reached the repository unreviewed.