Agents

MLCommons Agent Profile Named Hackathon Finalist

MLCommons' Agent Reliability Profile has reached the finals of a global hackathon, showcasing a new framework to ensure financial AI agents do not exceed their authorized boundaries.

ML Commons1 day agoAgents
Image: ML Commons

The MLCommons Financial Services Working Group has advanced to the final round of the C:>DIR Global "Agentic Regulator" Hackathon with its Agent Reliability Profile. Out of 336 submissions representing more than 65 countries, only 36 teams—six per problem space—made the cut. The competition, run by the University of Cambridge's Digital Innovation and Regulation Initiative, is backed by the BIS Innovation Hub, the Global Financial Innovation Network, the Digital Regulation Cooperation Forum, and over 35 other supporting organizations. The final-round builds took place from September 1 to September 8, with live demos and judging scheduled for September 15 and 16, ahead of the winner announcement on September 18, 2026, at the C:>DIR Summit.

Submitted by working group co-chairs Mike Hsu and Medha Bankhwal, the entry competes in the Know Your Agent, Digital Verification & Digital Public Infrastructure category. It showcases two distinct tools built on the Agent Reliability Profile framework. First, the Profile Builder ingests an institution's internal documentation, policies, and configuration files to generate a Level 1 Asserted Profile, highlighting any discrepancies between intended design and actual setup. Second, the Profile Validator uses this profile as a test specification, attempting to falsify it against live system behavior to produce a Level 2 Validated Profile.

For financial practitioners and regulators, this framework addresses what MLCommons describes as the "authorization and supervision gap." While current identity standards can verify an agent's identity, they fail to prove what actions that agent is permitted to take once integrated into financial systems. To demonstrate how the tools bridge this gap, the team applied them to an open-banking data-sharing and consent scenario, verifying whether an AI agent strictly adheres to the data-sharing limits authorized by a customer. This provides a standardized, regulator-compatible method for validating agentic deployments.

This is our own summary of reporting by ML Commons

More in Agents