Skip to main content
#89
Ranked #89 of 93 in this category· 该品类排名 #89 / 共 93 个

ai-evals

by RefoundAI·1mo ago

Help users create and run AI evaluations. Use when someone is building evals for LLM products, measuring model quality,

Testing & DebuggingClaude CodeCodexAutomated screen · 自动筛查open source · 开源

Before you install安装前须知

No special access needs declared未声明特殊权限需求
Editor's verdict· 编辑结论

Auto-published by GitHub discovery based on star threshold.

— Editorial team · 编辑团队

Install via Skills CLI

Use npx skills add to install this skill into the selected agent. Phase 0 commands are generated from source rules, not verified.

Codex
npx skills add https://github.com/RefoundAI/lenny-skills/tree/main/skills/ai-evals -g -a codex -y

Drop `-g` to install project-locally

Best for适合什么场景

  • High-star GitHub skill discovery candidate.

Not for不适合什么场景

  • Sensitive or production workflows without local review.

vs alternativesvs 其他选择

Full compare table完整对比表 →
#1QA Loop

Open the product, try the flow, fix what breaks, repeat.

4.9·15k stars
diff · 差异Best browser QA pick when you need evidence to leave a paper trail. Each run produces screenshots, console diffs, and a reproducible action log — much harder for stakeholders to wave off than "I tested it locally." Works well as a pre-merge gate and for filing bugs with repro steps attached. Not for unit tests, and not for authenticated production sessions where the screenshot itself becomes a data risk.
#2GStack Investigate

No fixes until the root cause is real.

4.8·123k stars
diff · 差异Best when the bug lives inside the code itself, not in operational state. Same "no fixes until the root cause is real" discipline as incident-investigate, but biased toward static code investigation: reads suspect modules, builds a hypothesis tree, asks for a failing test or repro before proposing a change. Strongest on flaky tests and intermittent failures where shallow patches make things worse. For ops-side incidents (logs, traffic, infra), incident-investigate fits better.
#3GStack QA

Open the app, test the flow, fix what breaks.

4.8·123k stars
diff · 差异Best when browser QA needs to close the loop — find the bug, propose the fix, verify the fix, leave evidence. Where qa-loop emphasizes evidence trails for stakeholder reporting, gstack-qa emphasizes shipping the fix in the same session. Strongest on frontend refactors and visual regressions. Same screenshot data-risk caveat as qa-loop: don't point it at authenticated production sessions where the screenshot itself becomes a leak.

Side-by-side compare维度对比

Key differences with same-lane alternatives
this skill · 当前ai-evalsQA LoopGStack InvestigateGStack QA
Rating · 评分4.94.84.8
Stars · 星标1.0k15k123k123k
Risk · 风险Automated screen · 自动筛查Medium risk · 中风险Low risk · 低风险Medium risk · 中风险
Best for · 最适合High-star GitHub skill discovery candidate.Browser smoke testsNo fixes until the root cause is real.Open the app, test the flow, fix what breaks.
Not for · 不适合Sensitive or production workflows without local review.Pure unit testingWorkflows that require stronger human review than this catalog entry documents.Workflows that require stronger human review than this catalog entry documents.

Audit notes审计备注

not individually audited · 未独立审计
Source源码open on GitHub · 公开
Author作者community · 社区!
Network网络访问not individually audited · 未独立审计!
Filesystem文件写入not individually audited · 未独立审计!
Dependencies依赖not individually audited · 未独立审计!
Telemetry遥测none · 无
Skill Market
Find the best AI skills for the job·按品类找最好用的 AI 技能
v0.4 · 1306 skills indexed · last review 2026-06-10