's   

Rushi's

Ctrl+AI+Ship

  • Home
  • Musings
  • Tech
  • About
  • Contact

Tag: llm-evaluation

May 05
2026
0

How to evaluate an AI skill

Posted by Rushi

You installed a skill. The README looks fine. The demo in the docs worked on the first try. Should you keep it? You can’t tell from the README. You have to run the skill on your own work and look at what comes out. This post walks through doing that. We start with the simplest […]

Read More →
tech agent-evaluation, ai-skills, ai-tooling, ai-workflow, anthropic-skills, claude, claude-agent-sdk, claude-cli, developer-tools, evals, llm-as-judge, llm-evaluation, prompt-engineering, skill-evaluation, skill-md, test-fixtures

Tags

ai AI agents AI coding agents angularjs anthropic artificial intelligence automation browser Chrome claude Claude Code code context window css cursor design developer productivity developer tools git Google html images java javascript js linux llm LLMs machine learning MCP nasa ollama open source productivity programming prompt engineering Python Research software engineering Spec-Driven Development typescript video videos Windows youtube

RSS RSS

  • Advent of Agents Season 3
  • Astryx: Meta’s Agent-Ready Design System, Reviewed
  • A History of Claude’s Published System Prompt
  • Playwright MCP vs CLI
  • Jev: A System One AI Decision Engine
October 2026
MTWTFSS
 1234
567891011
12131415161718
19202122232425
262728293031 
« Sep    
© 2026  rushis.com. | The content is copyrighted to Rushi and may not be reproduced.