✉news TechnologyAI first seen 19 h ago, last 19 h ago, peak #14
AI fellows warn labs run models with safeguards off
Original: 'We can't trust them completely': AI research fellows warn that labs are running models with the safeguards off behind closed doors
AI research fellows are warning that major AI laboratories may be testing frontier models in closed-door environments with safety guardrails disabled, meaning publicly demonstrated safeguards may not reflect how the systems actually behave during development. 'We can't trust them completely,' one fellow said, arguing that internal evaluations stripped of protections could hide risks from regulators and the public. The comments add to ongoing debate over transparency and oversight of advanced AI development.
Why now: Renewed concern over AI lab transparency and safety oversight is circulating in the news.
Evidence
API: https://socialmediatrends-api.osmike.com/v1/trends/792604