Unsupervised learning
Supervised learning works well when you already know what outcome you are trying to predict. But sometimes the business does not know what patterns exist yet. That is where unsupervised learning becomes useful.
Unsupervised learning looks through data and tries to find structure without being given the answer in advance. It can be useful when a business has a large amount of information and wants to explore it. Instead of asking the model to predict a specific outcome, we ask it to help us understand what is in the data.
This can be very useful in HSE and risk because the same issue is not always recorded in the same way. Two incidents might be closely related even if they have different classifications. Two hazards might have the same underlying cause even if they appear in different departments. Two sites might have similar risk profiles even if their headline metrics look different. Unsupervised learning can help reveal those patterns.
Clustering models
Clustering is one of the most common unsupervised learning techniques. It groups similar records together based on the information available.
A retailer might use clustering to group customers into segments. One group might be regular high-value customers. Another might be occasional bargain shoppers. Another might be people who only buy during promotions. The model is not told what groups to create. It works them out from the data.
For a mining HSE team, clustering can be a useful way to understand incidents, hazards or risk records in more detail. Take hand injuries as an example.
A normal report might show the total number of hand injuries by month, site or department. That is useful, but it may hide the fact that “hand injuries” is not really one problem. A clustering model might show that the data contains several different patterns:
- Maintenance-related hand injuries
- Tooling injuries
- Manual handling injuries
- Line-of-fire injuries
- Pinch-point injuries
- Injuries linked to isolation or stored energy
Clustering can help move the conversation from “we need to reduce hand injuries” to “we have three different hand injury problems, and each one needs a different control strategy.” That is much more actionable.
Topic modelling and text grouping
A lot of valuable business information is hidden in text fields. Incident descriptions, investigation findings, audit notes, risk assessments and corrective action comments often contain the detail that does not fit neatly into dropdown fields.
The problem is that text is hard to analyse at scale. If you have 50 incident reports, someone can read them. If you have 50,000, that is a very different story.
Topic modelling helps identify common themes across large volumes of text. For example, it might find that incident descriptions commonly refer to themes such as poor visibility, communication breakdowns, manual handling, line of fire, fatigue, contractor supervision or procedure compliance.
This can be powerful because the themes may not match the categories already used in the system. A business might have an incident classification called “slip, trip and fall”, but the written descriptions may reveal a recurring theme around poor lighting, uneven ground or housekeeping. Those details may be where the real improvement opportunity sits.
In an HSE setting, topic modelling can help answer questions like:
- What are people actually writing about in incident reports?
- Are new risk themes emerging?
- Are corrective actions addressing the same issues repeatedly?
- Do audit findings contain recurring language?
- Are similar hazards being described differently across sites?
This is one of the areas where machine learning can help make better use of data the business already has.
Association models
Association models look for things that often occur together. The classic example is supermarket basket analysis. If people often buy certain products together, the retailer can use that information for store layout, promotions or recommendations.
The same idea can be applied to risk. For example, a model might find that certain combinations appear together more often than expected:
- Vehicle interaction hazards and poor visibility
- Manual handling injuries and specific maintenance tasks
- Fatigue reports and particular roster patterns
- Environmental exceedances and certain weather conditions
- Overdue corrective actions and particular departments
- Contractor incidents and certain types of work
This can help teams see relationships that may not be obvious when looking at one field at a time. A safety report might show that vehicle interaction is a major hazard. An association model might show that vehicle interaction reports often include references to lighting, road conditions and radio communication. That gives the business a more useful starting point.
Instead of treating vehicle interaction as a broad category, leaders can begin asking more specific questions. Are the controls around visibility strong enough? Are road conditions changing? Are communication protocols being followed? Are the issues more common at certain times of day? Association models are useful because they help connect the dots.
Dimensionality reduction
Dimensionality reduction sounds technical, but the idea is fairly simple. Sometimes a dataset has too many variables to easily understand. A risk profile might include hundreds of measures across incidents, hazards, audits, training, inspections, workforce mix, environmental results, equipment, locations and controls. Looking at every measure one by one can become overwhelming.
Dimensionality reduction helps simplify the data while keeping the most important patterns. A simple way to think about it is compressing a complex picture into a smaller number of meaningful features. You lose some detail, but you keep the shape of what matters.
For HSE and risk teams, this can help with things like comparing sites, understanding overall risk profiles, or identifying the main drivers behind a complex set of indicators. For example, two sites may look different when you compare individual metrics. But after simplifying the data, they may turn out to have similar underlying risk patterns. Another site may look average on individual measures, but stand out when all indicators are considered together.
This type of analysis is often useful for exploration. It helps leaders see the bigger picture before drilling into the detail.
How these models can work together
In real projects, these techniques are often combined. A safety analytics solution might use clustering to group similar incidents, topic modelling to understand written descriptions, classification to flag high-risk records, and recommendation models to suggest relevant controls.
An operational analytics solution might use IoT data to detect abnormal equipment behaviour, forecast component failure, estimate remaining useful life, and optimise maintenance windows around production constraints.
The important thing is not the model type. The important thing is the business question.
A good machine learning project usually starts with questions like:
- What decision are we trying to improve?
- What are people currently doing manually?
- Where are teams spending time reviewing large volumes of data?
- What risks are difficult to see early?
- What information would help people act sooner?
- What patterns do we suspect exist, but cannot easily prove?
Once the question is clear, the model choice becomes much easier. If you need to predict a known outcome, supervised learning may be the right path. If you need to explore patterns in messy or complex data, unsupervised learning may be more useful. If you need to recommend an action, prioritise work or detect something unusual, there are specific model types that can support those goals.
What machine learning needs to work well
Machine learning does not need perfect data, but it does need useful data. That means the data needs to be available, relevant and consistent enough for the model to learn something meaningful.
It also needs a clear business purpose. “Use AI for safety” is too broad. “Identify incident reports that may have high-potential characteristics” is much clearer. “Improve risk management” is broad. “Group similar corrective actions so we can identify repeat issues” is something a model can actually help with. “Use IoT data” is broad. “Predict early signs of component failure so maintenance can be planned before unplanned downtime occurs” is much more useful.
The other important part is process. A prediction sitting in a dashboard does not create value by itself. The value comes when someone uses it to make a better decision, intervene earlier, prioritise work or ask a better question.
This matters even more in health, safety, environment and risk. These are not areas where businesses should blindly automate decisions. Human judgement, operational context and accountability still matter. Machine learning should support the decision-making process, not replace it.
A note on predicting safety incidents
There is one type of machine learning idea that often comes up in safety conversations, and it is worth addressing directly. It is the idea of training a model to predict when an accident or injury will occur.
On the surface, this sounds useful. Everyone wants fewer people to get hurt. If a model could predict the next incident, surely that would be valuable. The problem is that this framing can quickly become unhelpful.
Safety incidents are rare, complex and influenced by many human, operational and environmental factors. They are also high-consequence events. If the model gets it wrong, the cost is not a missed sales forecast or an inaccurate stock estimate. The cost may involve people’s health and safety.
That creates a “live or die by the sword” problem. If the model predicts an incident and nothing happens, people may lose trust in it. If the model does not predict an incident and something does happen, people may ask why the system failed. Either way, the conversation can become focused on whether the model predicted the event, rather than whether the business is managing risk better.
A more useful approach is to use machine learning to identify leading indicators, weak signals and patterns that deserve attention. For example:
- Which incident reports have high-potential characteristics?
- Which hazards are similar to previous serious events?
- Which controls are repeatedly failing verification?
- Which work areas are showing unusual reporting patterns?
- Which corrective actions are likely to become overdue?
- Which combinations of conditions are associated with higher risk?
- Which themes are emerging from investigation text?
These are better questions. They do not ask the model to predict the exact moment someone will be hurt. They ask the model to help the business see risk earlier, focus attention, and strengthen controls before harm occurs. That is a much better use of machine learning in safety.
Why machine learning matters
The real value of machine learning is not that it sounds impressive. The value is that it helps businesses use their data in ways that traditional reporting often cannot. Traditional reports are very good at telling us what happened. Machine learning can help us ask:
- What might happen next?
- What looks unusual?
- What patterns are we missing?
- Which records are similar?
- Which actions should we focus on?
- What information would help this person right now?
For mining businesses, this is especially relevant because teams are often dealing with large volumes of complex information. There may be thousands of incidents, hazards, inspections, audits, actions, monitoring results, maintenance records, equipment readings and risk records collected over many years. Inside that data are patterns that can help teams make better decisions.
Machine learning is one way to find those patterns. It will not replace expertise. It will not remove the need for good process. It will not magically fix poor data. But when it is applied to the right problem, with the right context around it, machine learning can help businesses move from reporting what has already happened to understanding what may need attention next. This is where it becomes useful.
In future articles, we will look at how machine learning algorithms are trained, why practices like training and test sets matter, and how different model types work in practice. We will also take a deeper look at individual techniques, such as classification, forecasting, recommendation systems, clustering and optimisation, with practical examples of how they can be used in real business settings.
Stick around for more articles on how these models work, where they are useful, and how to apply them in practice.
Wondering which of these fits a decision you are making?
That is exactly the conversation we like having. No jargon, no hard sell, just whether it is the right tool for the job.