About Grejna
Grejna AB is a Swedish research-based product company developing technology that makes uncertainty actionable in AI-assisted decisions. The name reflects the task itself: discerning when a model output is sufficiently well supported for use and when it requires review or further investigation.
Why Grejna?
Grejna is a newly coined Swedish name inspired by the Old Norse verb greina, a word associated with discerning, distinguishing and making things clear. The verb survives in modern Icelandic with related meanings including to distinguish, analyse and diagnose.
That idea captures what we are trying to achieve.
An AI model can produce an answer without that answer necessarily being sufficiently well supported for a real-world decision. Grejna helps organisations distinguish between outputs that can be used, cases that require human review and situations where the available evidence is not sufficient.
For us, to grejna means to systematically examine an AI output, make visible how well supported it is and what uncertainty remains, and provide a basis for deciding how it should be handled. Grejna does not make the organisation’s decisions.
Grejna: Discern when AI merits trust.
How Grejna works
Grejna is developing a decision-support layer for organisations that use predictive machine-learning models in real-world workflows. The solution complements the existing model with calibrated information about uncertainty, local explanations and a structured way to determine how each output should be handled.
The starting point is simple: a model can be accurate on average without being equally reliable in every individual case. For an organisation, it is therefore not enough to know the model’s average accuracy or receive a single probability value. It also needs to assess how well supported the individual output is, what influenced it and whether it can be used directly or should be reviewed by a person.
Grejna does not replace the customer’s model and does not make the organisation’s decisions. Instead, Grejna creates a controlled link between model outputs, uncertainty, explanations, decision rules, human handling and subsequent outcomes.
From average accuracy to the individual decision
Predictive models are usually evaluated using metrics that describe how the model performs across a larger body of data. These metrics are important, but they do not necessarily indicate how much an organisation should rely on an individual output.
Two cases can, for example, receive the same prediction even though the model has substantially stronger support for one than for the other. The model may have seen many similar cases in its training and calibration data, or it may be operating in a part of the data where the available evidence is limited. A high probability value does not necessarily mean that the probability is well calibrated either. If a model assigns an 80 per cent probability, the corresponding outcome should occur in roughly 80 per cent of comparable cases. Many models do not automatically satisfy this requirement.
This distinction matters when a model output is used to determine what happens next. If all outputs are treated the same, the organisation risks either automating cases that should have been reviewed or spending human time on cases where the model could have been used with sufficient confidence.
When we say that AI merits trust, we do not mean general trust in AI or that an individual output is guaranteed to be correct. We mean that a specific model output, under known conditions and for a defined use, has sufficient support according to the calibrated uncertainty information and the organisation’s established decision policy.
Research base
The research foundation: Calibrated Explanations
Grejna builds on the research behind Calibrated Explanations, developed by Helena Löfström, Tuwe Löfström, Ulf Johansson and Cecilia Sönströd.
Calibrated Explanations is a model-agnostic method for explaining individual predictions while making uncertainty explicit. The calibration mechanism depends on the type of predictive task.
A conventional local explanation method can show which variables had the greatest influence on a prediction. If the model’s original probabilities are poorly calibrated, however, the explanation can be difficult to use as the basis for a real decision. It is not enough for the explanation to represent the model accurately if the model itself provides a misleading picture of its own certainty.
For classification, the method can use Venn–Abers calibration to provide calibrated probability estimates together with information about uncertainty in those estimates. For regression, conformal methods are used to describe predictions, prediction intervals and probabilistic uncertainty. Calibration draws on a separate dataset with known outcomes.
Across these settings, the common objective is to complement the model output with calibrated uncertainty information and local explanations of what influenced the result. A narrower interval means that the estimate is more precisely bounded, while a wider interval indicates that the available evidence leaves greater room for uncertainty.
The uncertainty information is not a guarantee that an individual prediction is correct. It provides a statistical basis for distinguishing between more and less well-supported outputs.
The method
Explanations that also make uncertainty visible
Calibrated Explanations does not only explain the model’s final output. The method also describes how different features in the individual case influenced the calibrated probability.
Each feature contribution is given a defined meaning in relation to the model’s calibrated output. In simplified terms, the probability for the individual case is compared with the model’s estimate when a particular feature is changed in a systematic way. This makes it possible to see whether that feature increased or decreased the probability and how large its contribution was.
The feature contributions are also accompanied by uncertainty information. This means the user can see both the apparent importance of a feature and how confidently that importance can be estimated. This differs from explanations that present feature importance as exact numbers without showing how sensitive those values are to uncertainty in the model and the available evidence.
The method can provide two types of local explanations:
Factual explanations
A factual explanation describes the output produced by the model for the individual case. It shows the calibrated probability, the uncertainty interval and the features that contributed to the output.
The purpose is to give the user a clearer basis for understanding why the model made that assessment and how strongly supported it is.
Alternative explanations
An alternative, or counterfactual, explanation shows how the model’s assessment could change under different conditions. It can, for example, show what happens to the probability if the value of a particular variable is above or below a defined threshold.
These explanations show how the model responds to alternative inputs. They should not be interpreted as evidence that the change causes a particular real-world outcome or as a recommendation to the person affected by the decision.
Stability and research evaluation
A local explanation needs to be stable in order to be useful. If the same model, the same data and the same case produce different explanations across repeated runs, it becomes difficult to determine which explanation should form the basis for a decision.
Calibrated Explanations is designed so that the same case produces the same explanation as long as the model and calibration data remain unchanged. The published study evaluated the method’s stability and robustness across 25 datasets for binary classification. The results showed very high stability with an unchanged model and low variation when the model and calibration data were changed in a controlled way.
This does not mean that the explanation remains unchanged when the model or its environment genuinely changes. On the contrary, a correct explanation should reflect relevant changes in the model’s output. Stability here means that variation in the explanation should follow real changes rather than arise randomly between identical runs.
Grejna brings the research into an operational decision workflow
Calibrated Explanations provides technical evidence: a calibrated prediction, an uncertainty interval and a local explanation. The organisation then needs to determine what that information means for practical use.
Grejna is developing the layer that connects the research-based method to operational decisions. This layer is intended to answer questions that the explanation method itself is not designed to address:
- Is the output sufficiently well supported for the intended use?
- What level of risk has the organisation accepted?
- Should the output be used directly, reviewed by a person or escalated?
- Which model, calibration and decision rule were used?
- How did the human reviewer handle the case?
- What was the actual outcome?
- Can the entire decision process be reconstructed afterwards?
Grejna’s value lies in keeping these elements connected. Uncertainty should not merely be displayed in an interface. It should have a defined meaning for how the organisation acts.
Decision layer
How the decision layer works
1. The existing model produces an output
The process begins with the customer’s existing predictive model. Grejna is not intended to replace the model or the technical environment in which it is trained and used.
Instead, the model output is complemented with Calibrated Explanations. The customer can therefore retain the model and workflow already in place while gaining more useful evidence for assessing each output.
2. The output is calibrated and explained
Calibrated Explanations produces a calibrated probability, an uncertainty interval and a local explanation. Together, these describe what the model predicts, how precisely the probability can be estimated and which features influenced the output.
This information forms the technical decision evidence. It does not, however, determine on its own which action the organisation should take.
3. A decision rule translates the evidence into action
Grejna connects the technical evidence to a customer-specific decision policy. The policy defines the conditions that must be met for an output to be used directly and when a person needs to take over.
The decision policy can take into account factors such as:
- the model’s calibrated probability and uncertainty interval,
- the consequence of an incorrect decision,
- the type of case,
- the level of risk accepted by the organisation,
- the availability of human review,
- requirements for documentation or escalation.
Grejna does not determine what level of risk the customer should accept. The organisation owns the risk level and its decision rules. Grejna’s role is to make the relationship between technical evidence and operational action clear, consistent and possible to monitor.
Decision rules also need to be version-controlled. It must be possible to establish which rule was in force when a particular output was used, even if the organisation later changes its thresholds or processes.
4. The case is used, reviewed or escalated
Depending on the defined policy, a case can proceed in different ways. A sufficiently well-supported output can be used in the subsequent workflow. A more uncertain or higher-risk case can be sent for human review. A case that does not meet the defined conditions can be deferred or escalated.
Human oversight here means more than displaying an explanation on a screen. The right cases must reach the right person together with the relevant decision evidence. It must also be possible to document whether the reviewer followed the model output, changed the assessment or requested additional information.
5. The decision process is preserved
For each case, it must be possible to link the information that formed the basis for its handling. Depending on the customer’s requirements, the decision record can include:
- a reference to the individual case,
- the model and model version,
- the calibration dataset or calibration version,
- the model’s original output,
- the calibrated probability and uncertainty interval,
- the local explanation,
- the decision policy that was applied,
- the selected action,
- any human assessment,
- the time of the decision and the subsequent outcome.
The purpose is not to collect more information than necessary. It is to make it possible to reconstruct what the system and the organisation actually knew and did when the decision was made.
This is an important distinction from rerunning the same case afterwards using the latest version of the model. If the model, calibration or decision policy has changed, such a retrospective run can produce a different result from the one on which the original handling was based.
6. The outcome is linked back to the decision
When the actual outcome becomes known, it needs to be linked to the original decision. This makes it possible to assess how the model, uncertainty information and decision policy performed in practice.
The organisation can, for example, examine:
- how often automatically handled outputs were correct,
- whether the accepted risk level was maintained,
- which types of cases most often required human review,
- how often reviewers changed the model’s assessment,
- whether recurring types of cases require a different decision rule,
- whether the calibration still corresponds to real-world outcomes.
This closes the loop between prediction and real-world outcome. The resulting feedback can be used to reconsider decision rules, improve the workflow and, where necessary, update the calibration in a controlled way.
Trust needs to be monitored over time
A model and its operating environment are not static. Customer behaviour, processes, data sources and types of cases can change. The model itself can be updated and begin producing different probabilities. Calibration data that were previously representative may become less relevant over time.
It is therefore not enough to calibrate the model once and assume that the uncertainty information will always retain the same meaning.
Grejna’s product direction includes continuous assessment of whether the calibrated signal remains useful for the decision policy chosen by the organisation. The central question is not only whether the data or model have changed, but whether the organisation can still use the uncertainty information to distinguish between direct handling and human review.
If that relationship weakens, the organisation should be able to detect it and respond in a controlled way. This can mean temporarily sending more cases for review, adjusting the decision policy or recalibrating the model. Each change needs to be documented and version-controlled so that it is possible to establish when it was introduced and which decisions it affected.
Human oversight as part of the system
Human oversight only works when it is connected to clear responsibility and a real operational workflow. Simply allowing a user to open an explanation is not sufficient.
Grejna is intended to help the organisation define when human review is required, what evidence the reviewer should receive and how the human assessment should be documented. This makes it possible to focus limited review capacity on the cases where uncertainty or consequences justify it.
This can both reduce unnecessary manual review and make the review that is carried out more relevant. The objective is not maximum automation, but a balance in which automation, human judgement and risk can be monitored together.
Pilot
What we test in a pilot
A pilot begins with an existing predictive model and a defined decision workflow. Together with the customer, we identify which decision the model affects, what risk an incorrect output creates and which real-world outcome can be monitored.
We then examine whether suitable calibration data are available and how Calibrated Explanations can be connected to the model. We develop an initial decision policy and test how outputs can be divided between direct handling, human review and escalation.
The pilot can be carried out by replaying historical decisions or within a defined part of a live workflow. The baseline and target values are established together with the customer before the results are evaluated.
The central questions are:
- Can the uncertainty information distinguish between more and less well-supported outputs?
- Can suitable cases be handled without unnecessary manual review?
- Do directly handled cases remain within the level of risk defined by the customer?
- Does the reviewer receive better evidence in uncertain cases?
- Can decisions and human actions be reconstructed afterwards?
- Can the resulting outcomes be used for continued monitoring?
Metrics such as automation rate, proportion of manual review and review time therefore gain a clear context. A high automation rate is not an objective in itself. It is valuable only when the error rate simultaneously remains within the limit accepted by the customer.
A focused layer in the customer’s existing environment
Grejna is intended to work alongside the customer’s existing models, workflows and technical systems. The solution does not replace model development, case management, general operational monitoring or the organisation’s broader governance and compliance processes.
Grejna’s focused role is to connect:
- calibrated uncertainty and local explanations,
- the assessment of whether the signal remains useful,
- the organisation’s version-controlled decision rules,
- automated and human handling,
- decision evidence and traceability,
- subsequent outcomes and monitoring.
This can provide the organisation with stronger evidence for transparency, human oversight and documentation. Grejna does not, however, guarantee that every decision is correct or that an AI system automatically complies with every requirement of the EU AI Act.
From research to a commercial product
Calibrated Explanations forms the research-based core. Grejna is developing the commercial infrastructure required to use the method consistently within an organisation.
This means taking the technology from an individual explanation to a connected decision process: from model outputs and expressed uncertainty to decision policies, human handling, traceability and monitoring against real-world outcomes.
It is this connection between research and practical use that Grejna intends to validate together with its first pilot customers.
Would you like to explore whether Grejna creates value in your decision workflow?
We would be happy to have an initial discussion about your model, your use case and what would be relevant to test and measure in a pilot.