's   

Rushi's

Ctrl+AI+Ship

  • Home
  • Musings
  • Tech
  • About
  • Contact

Tag: llm-evaluation

May 05
2026
0

How to evaluate an AI skill

Posted by Rushi

You installed a skill. The README looks fine. The demo in the docs worked on the first try. Should you keep it? You can’t tell from the README. You have to run the skill on your own work and look at what comes out. This post walks through doing that. We start with the simplest […]

Read More →
tech agent-evaluation, ai-skills, ai-tooling, ai-workflow, anthropic-skills, claude, claude-agent-sdk, claude-cli, developer-tools, evals, llm-as-judge, llm-evaluation, prompt-engineering, skill-evaluation, skill-md, test-fixtures

Tags

ai AI agents AI coding agents angularjs anthropic artificial intelligence automation browser Chrome claude Claude Code code css cursor design developer tools git Google html images java javascript js linux llm LLMs machine learning MCP nasa node.js ollama open source pics productivity programming prompt engineering Python Research software engineering Spec-Driven Development typescript video videos Windows youtube

RSS RSS

  • From Prompting Agents to Loop Engineering
  • Local RAG Without the Storage Tax: A Hands-On Guide to LEANN
  • Building Your First Agent with Flue
  • The AI native playbook: people, agents, and context
  • Scanning agent skills before you trust them: a look at NVIDIA SkillSpector
August 2026
MTWTFSS
 12
3456789
10111213141516
17181920212223
24252627282930
31 
« Jun    
© 2026  rushis.com. | The content is copyrighted to Rushi and may not be reproduced.