Towards Machine-Readable Protocols for Responsible AI Agents
Abstract
Agents are increasingly relying on machine readable protocols such as model context protocol, agent-to-agent, skills, or data cards, which give agents a standardised interface to specialised tools and datasets to complete complex tasks. However, these protocols primarily prescribe when and how a specific tool, dataset, or model should be used, but do not surface responsible AI considerations, such as potential misuses or limitations, i.e., when tools or datasets should not be used or where their use might require additional guardrails. We observe that agents tend to ignore such limitations or do not follow responsible AI best practices and as a result, performance can suffer. To address this, we investigate a straightforward intervention: augmenting machine readable protocols with responsible AI information. Across multiple base models, domains, and tasks across the machine learning pipeline, we find that responsible AI augmentation reduces the severity of failures, broadens the exploration of accuracy-fairness trade-offs and can improve performance, but does not consistently change behaviour across all tasks. We further show that task-specific responsible AI prompting does not resolve (and may increase) these inconsistencies. Our work highlights the importance of considering responsible AI practices in the design of agentic frameworks and argues that machine readable protocols are a promising target for community investment given their scalability, durability, standardisation and auditability.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.