Dev4 min reading time

Anthropic Adds Plugin Evals to Claude Code: 6 Grader Types, a No-Plugin Baseline, and a CI Gate for Skills

MarkTechPost
Read full post
Anthropic introduced a new plugin evaluation workflow for Claude Code that tests plugins against realistic prompts, measures their impact by comparing runs with and without the plugin, and offers six grader types including automated and model-based scoring.

More in Dev

Dev24 min read

Monitoring production agent lifecycle with AWS DevOps Agent and AgentCore Evaluations

AWS Blog

Introducing the Agents API

Covered by 4 sources
Dev12 min read

Marketing ops as code: Automating events from planning to follow-up on GitHub

GitHub Blog (AI & ML)