How can we help you today?

Fill in the form below so we can explore ways to reach your goals or call us at 1800 577 346.

1 / 2
x
How can we help you?
One last step

Leave your details below and we'll be in touch.

Confirmation
2 / 2
x
Previous
Next step
Thanks! We have received your form submission, I'll get back to you shortly!
Oops! Something went wrong while submitting the form

OpenAI launches third-party assessment priorities with KPMG

By:
on
OpenAI launches third-party assessment priorities with KPMG
What's new: K-Startup Grand Challenge 2020 for Australian/New Zealand Startups! More information here.

You have probably sat through a vendor briefing in which a chief product officer assures you that their AI model is safe, fair and auditable. The slide deck is polished. The compliance language is careful. Then you ask who verified those claims and the room goes quiet. OpenAI has decided that silence is no longer acceptable. The company announced a set of priorities for third-party assessments of its models, naming KPMG as the first auditor to apply the framework. The move signals that frontier AI developers recognise they cannot grade their own homework indefinitely, even if the rubric they have published still leaves room for interpretation.

Third-party assessment is not a new idea in technology. Financial systems, medical devices and aviation software have long submitted to external review because the cost of failure is measured in lives and liability. AI has operated under a lighter regime. Models ship with technical documentation and internal red-teaming reports, but independent verification has been voluntary, inconsistent and often confined to academic partnerships that lack commercial accountability. OpenAI's framework attempts to formalise what good looks like: clear scope, qualified assessors and public disclosure of findings. The partnership with KPMG adds a recognisable name to the effort, one that enterprises already trust to audit their financial statements and operational controls.

What the priorities cover

The assessment priorities address three domains that matter to organisations deploying generative AI at scale. The first is model behaviour. Assessors will evaluate whether a model produces outputs that align with its stated purpose, whether it refuses harmful requests consistently and whether it exhibits biases that could undermine fairness in hiring, lending or customer service. The second domain is security. Third parties will probe for vulnerabilities that could allow prompt injection, data exfiltration or adversarial manipulation. The third is transparency. OpenAI expects assessors to verify that documentation is accurate, that limitations are disclosed and that users can understand how a model arrived at a particular output when the stakes are high.

KPMG's role is to apply these priorities to OpenAI's models and publish findings that other organisations can reference when making procurement decisions. The auditor brings expertise in risk management, control testing and regulatory compliance, but it does not bring deep technical fluency in machine learning. That gap is deliberate. OpenAI wants assessors who can translate model behaviour into business risk, not just count parameters or measure perplexity. The trade-off is that KPMG will rely on OpenAI's own tooling and access to conduct the review, which raises questions about independence that the framework does not fully resolve.

What this means for L&D and innovation leaders

Learning and development teams that are building AI literacy programmes should treat third-party assessments as a new category of evidence. Employees need to understand that a model's technical performance is not the same as its operational safety. A language model can score well on benchmarks and still produce outputs that violate company policy, expose sensitive data or reinforce stereotypes. Third-party reports offer a structured way to discuss those risks without requiring every learner to become a machine learning engineer. L&D leaders can use assessment findings to design scenarios that reflect real vulnerabilities, rather than hypothetical ones, and to show employees that even the most capable models have boundaries that matter in practice.

Innovation leaders face a different challenge. Third-party assessments create a paper trail that procurement teams and legal departments will demand before approving new AI tools. That trail is useful, but it is not sufficient. An assessment is a snapshot. It reflects the model's behaviour at a point in time, under conditions the assessor could test. It does not predict how the model will perform in your environment, with your data, under your usage patterns. Innovation teams need to build internal testing protocols that complement external assessments, not replace them. That means red-teaming your own use cases, monitoring outputs in production and maintaining the capability to audit model behaviour even when the vendor's report says everything is fine.

The OpenAI framework also exposes a gap that organisations should prepare to fill. Third-party assessments will proliferate as regulation tightens and enterprise buyers demand proof of due diligence. Not all assessments will be equal. Some will be rigorous, transparent and conducted by firms with relevant expertise. Others will be compliance theatre: expensive, superficial and designed to satisfy a checklist rather than surface real risk. L&D and innovation leaders need to develop the judgment to tell the difference. That judgment comes from understanding what good assessment looks like, which questions to ask about methodology and which red flags indicate that an auditor is more interested in billing hours than protecting your organisation.

The partnership with KPMG is a start, but it is not a solution. OpenAI has published priorities, not standards. The framework does not specify how often assessments should occur, what level of access assessors must receive or what happens when findings reveal a material risk. Those details will emerge as the market matures and as regulators impose requirements that voluntary frameworks cannot satisfy. In the meantime, organisations that are deploying AI at scale should treat third-party assessments as one input among many, not as a permission slip to stop asking hard questions.

Sources:

Workflow Podcast

The WorkFlow podcast is hosted by Steve Glaveski with a mission to help you unlock your potential to do more great work in far less time, whether you're working as part of a team or flying solo, and to set you up for a richer life.

No items found.
FREE EBOOK

100 DOS AND DON'TS FOR CORPORATE INNOVATION

To help you avoid stepping into these all too common pitfalls, we’ve reflected on our five years as an organization working on corporate innovation programs across the globe, and have prepared 100 DOs and DON’Ts.

No items found.
No items found.

STEP INTO THE METAVERSE

Unlock new opportunities and markets by taking your brand into the brave new world.

Thanks for your submission. We will be in touch shortly!
Oops! Something went wrong while submitting the form.

Ask me a question!
THE ULTIMATE GUIDE TO
 AI IN REAL ESTATE
FREE EBOOK
Thank you! Your submission has been received!
Oops! Something went wrong while submitting the form.
RELATED
Get the latest content first
Thanks! We'll get back to you shortly!
Oops! Something went wrong while submitting the form
By signing up you agree to Collective Campus' Terms.