The 2026 revision of the Artificial Intelligence standard (GAI-01) has now been published, and the most significant change is one of emphasis rather than scope. The six capability areas are unchanged. What has moved is where assessors spend their time: evaluation and human oversight now carry considerably more weight in the portfolio and the professional discussion than they did in the previous edition.
Why the council made the change
The AI employer council reviewed a year of portfolio submissions alongside feedback from the organisations that hire our holders. The pattern was consistent. Candidates were strong on building models and noticeably weaker on showing how they knew a model was fit to ship. Benchmark scores were plentiful; evidence that a test reflected real use was not.
A model that performs well on a favourable benchmark tells you very little. We want to see the test that could have failed, and what the candidate did when it did.
Employer panels were equally clear that the costliest failures they had seen were rarely about architecture. They were about models that behaved differently in production than in testing, and about nobody having a plan for when that happened.
What assessors now look for
From the next assessment window, portfolios at every level will be expected to show evaluation evidence proportionate to the role. In practice that means:
- At BIPS Practitioner, a test design for a defined problem, with a clear explanation of what it does and does not cover, reviewed by a senior colleague.
- At BIPS Professional, evidence that you owned evaluation for a feature in production, including monitoring and at least one decision driven by what monitoring showed.
- At BIPS Specialist and above, evaluation practice you set for others, such as red-teaming protocols, release gates or safety reviews, with peer-reviewed evidence of its effect.
- At every level, a described escalation path for when the model is wrong, and a named point at which a person takes the decision back.
The evidence review stage will also check provenance more closely for evaluation work. Assessors will want to see that the tests were designed by the candidate, not simply inherited from a template or a library default.
What it means for current candidates
Candidates who have already submitted a portfolio will be assessed against the edition in force when they submitted. Anyone preparing a portfolio now should build evaluation and oversight evidence in from the start rather than treating it as an appendix. The professional discussion will routinely include questions about a test that surprised you, and what you changed as a result.
Existing BIPS-P, BIPS-Pro and BIPS-S holders are not required to reassess. The revised expectations will, however, inform CPD guidance for the coming year, and we recommend holders log evaluation-focused development against their annual hours.
The full revised descriptors are available from the standard page, and our advisers can talk through how the changes apply at your level.