AI News
Anthropic Puts Independent Evaluators Inside the Lab
Anthropic and Accenture plan to embed independent evaluators inside frontier-model development. Access, funding and release authority will determine whether it works.
AI News
Anthropic and Accenture plan to embed independent evaluators inside frontier-model development. Access, funding and release authority will determine whether it works.
AI News
Gemini now connects to work, creative and lifestyle apps. The product challenge is making every delegated action understandable and recoverable.
AI News
OpenAI added monitoring and controls for GPT-6 prompt caching. Agent teams now need to manage cache hit rate as a product metric.
AI News
Salesforce built Koa by training an open model on synthetic CRM workflows. The moat may be the workflow specification, not the base model.
AI News
Google’s new anomaly detector audits how enterprise agents reason and use tools. Task completion alone is no longer enough.
AI News
OpenAI is testing Sponsored Agents inside ChatGPT ads. PMs should evaluate the full question-to-handoff journey, not just click-through rate.
AI News
NVIDIA agreed to acquire Hugging Face for $12.93 billion. Teams should now test whether the platform remains neutral in practice.
AI News
GitHub can now hand up to 25 code-quality findings to Copilot. The useful metric is accepted fixes, not findings closed.
AI News
Images 2.5 improves targeted edits, reference fidelity and iteration speed. The product metric should be time to an approved asset.
AI News
The OpenAI–Hugging Face incident shows why agent evaluations need production-grade access controls, monitoring, stop paths, and recovery.
AI News
Gemini 3.6 Flash may lower token spend. PMs should compare cost per accepted task across model calls, tools, review, and recovery.
AI News
OpenAI Presence packages the operating layer around enterprise agents. PMs still need clear authority, handoffs, ownership, and accepted-outcome metrics.