How Skill Market decides what belongs here.Skill Market 的评测方法
We prefer small, explainable recommendations over a noisy index. Every featured skill should have a clear scenario fit, install path, and visible risk note.last reviewed · 上次评测: 2026-06-10
Four dimensions, weights public4 个维度,权重公开
Scoring formula评分公式Each candidate scores 0–10 on four dimensions, weighted as below into a single total. The top 3 per category enter "Editor's picks"; everything else stays searchable as "Community indexed" but doesn't appear in the default list.
Blind-tested against real task sets. We hide skill names, run the same scenarios twice, and measure whether the output actually solves the problem. We score scenario hit-rate, not general chat quality.
Public source, verified author, minimal network/filesystem/dependency surface, audit notes attached. Anything scoring below 5 on this dimension is blocked from "featured" regardless of other scores.
Commit activity in the last 90 days, issue response latency, release cadence, doc update frequency. Three months without maintenance, or three unanswered open issues — automatic drop from featured.
Install command, prerequisites, and platform support written on the card. Time-to-first-success measured under 5 minutes. "Try and find out" doesn't make the list.
When we re-rank, when we hot-fix什么时候重排,什么时候紧急复查
Review cadence评测节奏End of each quarter, the #1 and runner-up in every category go through a fresh blind test. If the #1 loses, ranks swap directly and the changelog records the reason.
CVE, leaked secrets, trust-boundary regressions — the affected skill is reviewed within 7 days. The review can demote, attach a red warning, or remove from featured entirely.
The GitHub crawler runs weekly. New candidates enter as "Community indexed" with a Curator review pending tag. Nothing gets featured without a manual review pass — full stop.
Code-review category: why these three以 code-review 为例:为什么是这 3 个
Walkthrough案例走查We picked code-review for the walkthrough because its Top 3 represent three different mental models for the same job — the formula weights show their shape most clearly there.
Scenario 9.4 / Security 8.8 / Maintenance 8.5 / Install 9.0 → 8.97. Strongest on diff-oriented concreteness and trust-boundary catches (SQL injection, auth, conditional side-effects). Doesn't re-summarize the repo — reads what changed.
Scenario 8.6 / Security 9.1 / Maintenance 8.0 / Install 8.8 → 8.55. Lighter than #1 — suggestion-only, no diff edits, no comment threads. Best for senior engineers second-opinion-ing their own PR before merge; doesn't compete with #1's "gate others' code" role.
Scenario 8.0 / Security 8.5 / Maintenance 7.8 / Install 9.2 → 8.32. Fills the narrow gap between "branch is ready" and "reviewer can pick it up cold" — doesn't audit logic, focuses on pre-flight checks and reviewer handoff notes.
Small, transparent, not pretending to be "we"小而透明,不假装是「我们」
Who reviews编辑团队Skill Market is currently run by one editor plus a network of review consultants. That's why featured counts are capped (currently 28 featured total, averaging 2-3 per category) — and why we changed the Hero copy from "Top 10 per category" to "Top picks per category": don't lie, prefer fewer entries. Every review decision lands in the changelog. Disagree with a verdict? Open an issue on the GitHub repo.
The last four re-rank decisions最近 4 次评测决策的来龙去脉
Re-rank changelog重排日志Previously each featured skill's editor verdict was a single sentence — not enough to support a real decision. Expanded the 14 featured we have full context on (gstack-review, qa-loop, etc.) to 80-120 words each in a consistent shape: positioning / strongest case / weakest case / small caveat.
The original "Top 10 per category" promise had only 2.8 featured per category in reality. Downgraded to "Top picks per category" to match the data; /skills also now defaults to showing only the 28 curated entries, with a chip to switch to the full 1229 indexed view.
In the new blind test round, the old #3 and #4 lost ground on the eval-discipline dimension and no longer qualify for featured. They remain searchable in Community indexed but no longer appear on the category first screen.
The two skills overlap little on the "evidence delivery vs. implementation fix" mental models. Lumping them together at one rank was misleading. After the split, qa-loop owns "evidence trail for stakeholders" and gstack-qa owns "fix in the same session".
Why our rankings are trustworthy为什么我们的排名值得信
Curation principles评测原则Pick the best fit for the actual scenario over the popularity contest. Wider is rarely better.
Install command, prerequisites, and platform support live on the card — no "try and find out".
Risk level, source link, and audit notes sit next to the install button, not on a separate page.
Each language page declares its own translation readiness; gaps are explicit, not hidden.
High-risk skills can be listed, but only with explicit red warnings and constraints.