Skip to main content
Methodology评测方法

How Skill Market decides what belongs here.Skill Market 的评测方法

We prefer small, explainable recommendations over a noisy index. Every featured skill should have a clear scenario fit, install path, and visible risk note.last reviewed · 上次评测: 2026-06-10

Four dimensions, weights public4 个维度,权重公开

Scoring formula评分公式

Each candidate scores 0–10 on four dimensions, weighted as below into a single total. The top 3 per category enter "Editor's picks"; everything else stays searchable as "Community indexed" but doesn't appear in the default list.

Scenario fit · 场景适配度35%

Blind-tested against real task sets. We hide skill names, run the same scenarios twice, and measure whether the output actually solves the problem. We score scenario hit-rate, not general chat quality.

Security & trust · 安全 & 信任25%

Public source, verified author, minimal network/filesystem/dependency surface, audit notes attached. Anything scoring below 5 on this dimension is blocked from "featured" regardless of other scores.

Maintenance signal · 维护信号20%

Commit activity in the last 90 days, issue response latency, release cadence, doc update frequency. Three months without maintenance, or three unanswered open issues — automatic drop from featured.

Install clarity · 安装清晰度20%

Install command, prerequisites, and platform support written on the card. Time-to-first-success measured under 5 minutes. "Try and find out" doesn't make the list.

When we re-rank, when we hot-fix什么时候重排,什么时候紧急复查

Review cadence评测节奏
01
Quarterly re-rank · 季度重排

End of each quarter, the #1 and runner-up in every category go through a fresh blind test. If the #1 loses, ranks swap directly and the changelog records the reason.

02
Security hot-fix · 安全紧急复查

CVE, leaked secrets, trust-boundary regressions — the affected skill is reviewed within 7 days. The review can demote, attach a red warning, or remove from featured entirely.

03
Community auto-discovery · 社区自动发现

The GitHub crawler runs weekly. New candidates enter as "Community indexed" with a Curator review pending tag. Nothing gets featured without a manual review pass — full stop.

Code-review category: why these three以 code-review 为例:为什么是这 3 个

Walkthrough案例走查

We picked code-review for the walkthrough because its Top 3 represent three different mental models for the same job — the formula weights show their shape most clearly there.

Small, transparent, not pretending to be "we"小而透明,不假装是「我们」

Who reviews编辑团队

Skill Market is currently run by one editor plus a network of review consultants. That's why featured counts are capped (currently 28 featured total, averaging 2-3 per category) — and why we changed the Hero copy from "Top 10 per category" to "Top picks per category": don't lie, prefer fewer entries. Every review decision lands in the changelog. Disagree with a verdict? Open an issue on the GitHub repo.

The last four re-rank decisions最近 4 次评测决策的来龙去脉

Re-rank changelog重排日志
2026-05-19
Editor's Verdict expanded to ~100 words for 14 featured

Previously each featured skill's editor verdict was a single sentence — not enough to support a real decision. Expanded the 14 featured we have full context on (gstack-review, qa-loop, etc.) to 80-120 words each in a consistent shape: positioning / strongest case / weakest case / small caveat.

2026-05-18
Hero copy Top 10 → Top picks

The original "Top 10 per category" promise had only 2.8 featured per category in reality. Downgraded to "Top picks per category" to match the data; /skills also now defaults to showing only the 28 curated entries, with a chip to switch to the full 1229 indexed view.

2026-04
agent-building featured cut from 4 to 2

In the new blind test round, the old #3 and #4 lost ground on the eval-discipline dimension and no longer qualify for featured. They remain searchable in Community indexed but no longer appear on the category first screen.

2026-03
testing split into qa-loop / gstack-qa

The two skills overlap little on the "evidence delivery vs. implementation fix" mental models. Lumping them together at one rank was misleading. After the split, qa-loop owns "evidence trail for stakeholders" and gstack-qa owns "fix in the same session".

Why our rankings are trustworthy为什么我们的排名值得信

Curation principles评测原则
01
Scenario fit beats raw popularity.

Pick the best fit for the actual scenario over the popularity contest. Wider is rarely better.

02
Install clarity is part of quality.

Install command, prerequisites, and platform support live on the card — no "try and find out".

03
Security notes must be close to the install decision.

Risk level, source link, and audit notes sit next to the install button, not on a separate page.

04
Chinese and English pages should disclose translation readiness.

Each language page declares its own translation readiness; gaps are explicit, not hidden.

05
High-risk skills can be listed, but only with explicit warnings.

High-risk skills can be listed, but only with explicit red warnings and constraints.

Skill Market
Find the best AI skills for the job·按品类找最好用的 AI 技能
v0.4 · 1306 skills indexed · last review 2026-06-10