Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
8 changes: 5 additions & 3 deletions AGENTS.md
Original file line number Diff line number Diff line change
Expand Up @@ -66,7 +66,7 @@ app/
twitter-image.tsx # next/og convention — twitter:image, re-exports opengraph-image (auto-wired)
og/repo/[...slug]/route.tsx # per-repo OG image — a plain route, because the next/og file
# convention would sit at a path the catch-all page route also claims
score/page.tsx # Live Score entry — URL form, past scores, FAQ
score/page.tsx # Live Score entry — URL form, visitor's recent scores + top leaderboard repos, FAQ
score/opengraph-image.tsx # next/og convention — Live Score OG image (auto-wired)
score/twitter-image.tsx # next/og convention — /score twitter:image, re-exports (auto-wired)
score/[host]/[owner]/[name]/page.tsx # live score; result cached 1h per repo (unstable_cache); redirects to the repo page when indexed
Expand All @@ -81,7 +81,7 @@ components/ # Tailwind-styled React components
HostPill.tsx, HostSelect.tsx, Medal.tsx, ModelPills.tsx,
MobileNav.tsx, Pagination.tsx, SearchBar.tsx, SelectMenu.tsx, SortSelect.tsx,
SignalRow.tsx, SuggestionItem.tsx, VersionPill.tsx,
RepoHero.tsx, ScoreDeltaPopover.tsx, SignalListCard.tsx, ModelSuggestions.tsx, PerModelScores.tsx,
RepoHero.tsx, RepoSummary.tsx, ScoreDeltaPopover.tsx, SignalListCard.tsx, ModelSuggestions.tsx, PerModelScores.tsx,
AlternativesStrip.tsx, BreadcrumbJsonLd.tsx, HomeJsonLd.tsx, ExternalLink.tsx,
BadgeEmbed.tsx, ActionEmbed.tsx, PeerlistCard.tsx, PeerlistBadge.tsx, ProductHuntBadge.tsx,
CopySnippet.tsx, PackageLookupForm.tsx,
Expand All @@ -99,6 +99,7 @@ lib/
badge.ts # SVG badge renderer (used by /api/badge)
contact.ts # packageRequestIssueUrl — pre-filled GitHub issue link for unscored packages
repo-path.ts # repoPath / repoIdentity — the one place a repo URL is built or parsed
repo-summary.ts # per-repo rank, best/worst agent, key signals — the repo-specific text on repo pages
scoring/
signals/ # one file per signal + helpers + types + index
weights.ts # per-model weight tables
Expand Down Expand Up @@ -133,6 +134,7 @@ tests/
parse-repo-url.test.ts # GH / GL / BB parsing + edge cases
scorer.test.ts # scoreRepo, topImprovements
badge-adoption.test.ts # detectBadgeEmbed — README badge-embed detection
repo-summary.test.ts # summarizeRepo / keySignals — ranks, ties, best/worst agent
path-resolution.test.ts # firstExisting / resolveRelative / resolveAllRelative — case-insensitive lookup
live-score.test.ts # content-candidate coverage vs the signals, path traversal, host URLs
signals/ # one *.test.ts per signal
Expand Down Expand Up @@ -210,7 +212,7 @@ If either sibling isn't present locally, flag it; never silently skip the propag

1. Add a `ModelProfile` to `MODELS` in `lib/scoring/weights.ts` — weights for every signal.
2. Appears automatically in the leaderboard model pills, methodology weight-profile panel, and repo-page suggestions.
3. Update the hard-coded agent lists/counts that **don't** derive from `MODELS`: `APP_DESCRIPTION` + `APP_KEYWORDS` in `lib/version.ts`, `lib/skill-content.ts`, the "Which agents" FAQ in `app/methodology/page.tsx`, `app/page.tsx`, `app/skill/page.tsx`, `app/action/page.tsx`, `app/terms/page.tsx`, `app/opengraph-image.tsx`, the footer strip in `app/og/repo/[...slug]/route.tsx`, the `generateMetadata` description + JSON-LD description in `app/repo/[...slug]/page.tsx`, the Dataset description in `components/HomeJsonLd.tsx`, and `README.md`. Grep the current count word (e.g. `eight`) and the trailing `OpenHands, Pi` to find them all — `tasks/` and `lib/changelog.ts` are historical records and stay as-shipped, and `.claude/skills/agent-friendly/SKILL.md` is an install artifact pinned by `skills-lock.json` (it refreshes via `npx skills add`, never by hand).
3. Update the hard-coded agent lists/counts that **don't** derive from `MODELS`: `APP_DESCRIPTION` + `APP_KEYWORDS` in `lib/version.ts`, `lib/skill-content.ts`, the "Which agents" FAQ in `app/methodology/page.tsx`, `app/page.tsx`, `app/score/page.tsx`, `app/skill/page.tsx`, `app/action/page.tsx`, `app/terms/page.tsx`, `app/opengraph-image.tsx`, the footer strip in `app/og/repo/[...slug]/route.tsx`, the Dataset description in `components/HomeJsonLd.tsx`, and `README.md`. Grep the current count as a word and a digit (e.g. `nine`, `9`) and the trailing `OpenHands, Pi` to find them all — `tasks/` and `lib/changelog.ts` are historical records and stay as-shipped, and `.claude/skills/agent-friendly/SKILL.md` is an install artifact pinned by `skills-lock.json` (it refreshes via `npx skills add`, never by hand).
4. Existing repos keep their old per-model rows until rescored. `PerModelScores` renders the missing model as "—" (not 0), but **everything that reads `model_score` via a JOIN degrades silently to empty** until the backfill lands: the leaderboard (`listLeaderboard`, i.e. `/?model=<new-id>`) and `getAlternatives` (the repo page's alternatives strip under `?model=<new-id>`) both inner-join `model_score` and return **zero rows** for a model with no rows yet — no error, just an empty page. Reweighting an existing model has the same staleness problem in reverse: stored scores stay at the old weights until rescored. Either trigger `scheduled-rescore.yml` via `workflow_dispatch` at deploy time, or accept up to 6 hours of the empty/stale state until the cron runs.
5. **Mirror to both siblings**: copy the weights change into `../agent-friendly-action/src/scoring/weights.ts` **and** `../agent-friendly-skill/src/scoring/weights.ts`, and log under "Unreleased" in each sibling's `CHANGELOG.md`.

Expand Down
4 changes: 2 additions & 2 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -193,7 +193,7 @@ GitLab and Bitbucket are implemented and score identically to a clone, but ship

## Companion: PR-diff GitHub Action

[`hsnice16/agent-friendly-action`](https://github.com/hsnice16/agent-friendly-action) runs the same scorer inside your CI and posts a per-PR score-delta comment — _"this PR drops your Claude Code score by 4.1 points because it removed CI config."_ Opt-in via an `AGENTS_BADGE_TOKEN` secret; falls through silently when unset. Each repo detail page on the dashboard ships a copy-paste workflow snippet under "Catch score regressions on every PR".
[`hsnice16/agent-friendly-action`](https://github.com/hsnice16/agent-friendly-action) runs the same scorer inside your CI and posts a per-PR score-delta comment — _"this PR drops your Claude Code score by 4.1 points because it removed CI config."_ Opt-in via an `AGENTS_BADGE_TOKEN` secret; falls through silently when unset. Each repo detail page on the dashboard ships a copy-paste workflow snippet under "Check the score on every pull request".

## Companion: agent skill

Expand All @@ -204,7 +204,7 @@ GitLab and Bitbucket are implemented and score identically to a clone, but ship
Read-only JSON endpoints for external integrators (skills, hooks, browser overlays, third-party tools):

- `GET /api/score?host=<host>&repo=<owner>/<name>` — look up an indexed repo by host + owner/name. Returns `{ repo, signals, modelScores }` on 200; `{ error: "not_indexed" }` with status 404 when the repo isn't in our DB. The natural lookup endpoint for any tool that has a repo URL but not our internal id.
- `GET /api/repos` — full leaderboard (id, owner, name, host, stars, overall_score, per-model scores).
- `GET /api/repos` — full leaderboard (each repo row plus its per-model scores).
- `GET /api/repo/<id>` — per-repo detail (signals, model scores, top improvements). Requires the internal id; use `/api/score` first if you only have host + owner/name.
- `GET /api/badge/<host>/<owner>/<name>.svg` — embeddable SVG badge. `?model=<id>` for per-model variants.
- `GET /api/package/<registry>/<name>` — resolve npm / PyPI / Cargo package → source-repo score (or `unresolved` when the registry doesn't expose a repo URL).
Expand Down
48 changes: 23 additions & 25 deletions app/about/page.tsx
Original file line number Diff line number Diff line change
Expand Up @@ -17,7 +17,7 @@ export const metadata: Metadata = {
url: "/about",
type: "article",
},
description: `Who built ${APP_NAME}, why it exists, and what it isn't. Independent, MIT-licensed, no affiliation with any AI agent vendor.`,
description: `Who built ${APP_NAME}, why, and what it is not. Independent, MIT-licensed, and not tied to any AI agent company.`,
};

const ABOUT_JSON_LD = {
Expand Down Expand Up @@ -59,50 +59,47 @@ export default function AboutPage() {

<section className="my-3 mb-7">
<h1 className="mb-2.5 text-[30px] font-bold leading-[1.18] tracking-tight">About</h1>
<p className="m-0 max-w-[72ch] text-[15.5px] text-ink-dim">
Who built {APP_NAME}, why it exists, and what it deliberately isn&apos;t.
</p>
<p className="m-0 max-w-[72ch] text-[15.5px] text-ink-dim">Who built {APP_NAME}, why, and what it is not.</p>
</section>

<Panel>
<PanelHeading>Who</PanelHeading>
<p className="m-0 text-[14.5px] leading-relaxed text-ink-dim">
Built and maintained by <ExternalLink href="https://github.com/hsnice16">Himanshu Singh</ExternalLink>.
Independent project — no affiliation with Anthropic, OpenAI, Google, Cognition, Anysphere, or any of the agent
vendors ranked here.
Built and maintained by <ExternalLink href="https://github.com/hsnice16">Himanshu Singh</ExternalLink>. It is
an independent project, not tied to Anthropic, OpenAI, Google, Cognition, Anysphere, or any other company
whose agent is ranked here.
</p>
</Panel>

<div className="mt-3.5">
<Panel tone="warn">
<PanelHeading tone="warn">Why this exists</PanelHeading>
<PanelHeading tone="warn">Why it exists</PanelHeading>
<p className="m-0 text-[14.5px] leading-relaxed text-ink-dim">
The gap between &ldquo;repo with a README&rdquo; and &ldquo;repo that actually helps an AI coding agent ship
code&rdquo; keeps widening, and there&apos;s no public way to tell who&apos;s doing the work. {APP_NAME}{" "}
tries to make that visible — per model, because the agents aren&apos;t interchangeable. Claude Code wants an
AGENTS.md and a fast test loop; Cursor wants strong types and a skim-readable README; Devin wants a runnable
dev environment with declared deps and tests. The same repository can score very differently across them,
and a single overall number would hide that.
A repo that has a README is not the same as a repo that really helps an AI coding agent get work done. That
gap keeps growing, and there is no public way to see which repos have done the work. {APP_NAME} tries to
show it, for each agent separately, because the agents are not the same. Claude Code wants an AGENTS.md and
fast tests. Cursor wants strong types and a README that is easy to skim. Devin wants a dev setup it can run,
with its dependencies and tests listed. One repo can score very differently for each of them. A single
number would hide that.
</p>
</Panel>
</div>

<div className="mt-3.5">
<Panel>
<PanelHeading>What it isn&apos;t</PanelHeading>
<PanelHeading>What it is not</PanelHeading>
<p className="m-0 text-[14.5px] leading-relaxed text-ink-dim">
This is not a benchmark of agent performance. Today every score is derived from{" "}
<strong className="text-ink">static signals</strong> — file existence and content-length checks on the
cloned tree. No agent is actually run. Per-model rationales are derived from each agent&apos;s published
documentation (sources are linked on the methodology page), but the weight values themselves are still
pre-benchmark — not yet calibrated against measured agent success. Read the{" "}
It is not a test of how well agents perform. Every score comes from{" "}
<strong className="text-ink">simple file checks</strong>: does a file exist, and how long is it. No agent is
actually run. The reasons behind each agent&apos;s weights come from that agent&apos;s own docs (linked on
the methodology page). But the weight numbers are not yet tested against how agents really perform. Read the{" "}
<Link
href="/methodology"
className="border-b border-dotted border-ink-dim/60 text-ink-dim hover:border-ink-soft hover:text-ink-soft"
>
methodology
</Link>{" "}
for the full picture, including the production-cut plan to replace pre-benchmark weights with measured ones.
for the details, including the plan to replace these weights with measured ones.
</p>
</Panel>
</div>
Expand All @@ -111,9 +108,9 @@ export default function AboutPage() {
<Panel>
<PanelHeading>Open source</PanelHeading>
<p className="m-0 text-[14.5px] leading-relaxed text-ink-dim">
MIT-licensed. The signal definitions, weight profiles, scoring code, seed list, and every score in the
database are all in the <ExternalLink href={REPO_URL}>source repository</ExternalLink>. If a repo&apos;s
score looks wrong, file an issue with a link and the rubric to revisit; if a signal is missing, propose one.
MIT-licensed. The checks, the weights, the scoring code, the list of repos, and every score are all in the{" "}
<ExternalLink href={REPO_URL}>source repository</ExternalLink>. If a score looks wrong, open an issue with a
link and say which rule to look at again. If a check is missing, suggest one.
</p>
</Panel>
</div>
Expand All @@ -123,7 +120,8 @@ export default function AboutPage() {
<PanelHeading>Contact</PanelHeading>

<p className="m-0 text-[14.5px] leading-relaxed text-ink-dim">
Best signal: open an issue or discussion on <ExternalLink href={`${REPO_URL}/issues`}>GitHub</ExternalLink>.
The best way to reach me: open an issue or discussion on{" "}
<ExternalLink href={`${REPO_URL}/issues`}>GitHub</ExternalLink>.
</p>
</Panel>
</div>
Expand Down
Loading
Loading