feat(skills): 新增 gitlink-newcomer 新人引导 Skill #125
|
|
@ -0,0 +1,112 @@
|
|||
---
|
||||
name: gitlink-newcomer
|
||||
version: 1.0.0
|
||||
description: "新人引导:识别 good-first-issue、评估上手难度与友好度、为候选 Issue 生成个性化引导评论、产出新手任务看板,降低新贡献者参与门槛。当用户提到「适合新人的 Issue」「good first issue」「新手任务」「引导新贡献者」「降低参与门槛」「new contributor」时触发。"
|
||||
metadata:
|
||||
requires:
|
||||
bins: ["gitlink-cli"]
|
||||
optional_bins: ["python"]
|
||||
cliHelp: "gitlink-cli issue --help"
|
||||
---
|
||||
|
||||
# gitlink-newcomer(新人引导)
|
||||
|
||||
**CRITICAL — 开始前先阅读 [`../gitlink-shared/SKILL.md`](../gitlink-shared/SKILL.md),其中包含认证、权限处理和 API 注意事项。**
|
||||
**CRITICAL — GitLink 操作只能用 `gitlink-cli`。禁止用 `gh`(GitHub CLI)操作 GitLink 资源。`gh` 仅适用于 GitHub 平台。**
|
||||
**CRITICAL — 发布引导评论属于写操作,执行前必须征得用户确认。本技能默认只生成评论内容,不自动发布。**
|
||||
|
||||
> **前置条件:** 先阅读 [`../gitlink-shared/SKILL.md`](../gitlink-shared/SKILL.md) 了解认证与全局参数。
|
||||
|
||||
## 何时使用本技能
|
||||
|
||||
- 维护者想为新贡献者整理一份「适合上手的 Issue」清单
|
||||
- 用户问「这个项目有哪些 good first issue / 新手任务」
|
||||
- 想为某个简单 Issue 自动生成欢迎与引导评论,降低新人门槛
|
||||
- 需要评估某个 Issue 对新手的友好程度
|
||||
|
||||
## 何时不使用
|
||||
|
||||
- 只是普通地列出全部 Issue → 用 `gitlink-issue`
|
||||
- 复杂的项目健康度/协作分析 → 用 `gitlink-health` / `gitlink-insight`
|
||||
|
||||
## 能力概览
|
||||
|
||||
| 能力 | 说明 |
|
||||
|------|------|
|
||||
| good-first-issue 识别 | 综合标签、标题/正文关键词、描述长度、讨论热度等多信号识别 |
|
||||
| 友好度评分 | 为每个 Issue 输出 0-100 的新手友好度分与难度等级(入门/较易/中等/进阶) |
|
||||
| 个性化引导 | 为候选 Issue 生成包含上手步骤、Fork/PR 流程的欢迎评论 |
|
||||
| 新手任务看板 | 汇总候选 Issue 为 Markdown 看板,可贴到 Wiki 或 README |
|
||||
|
||||
## 工作流 1:生成新手任务看板
|
||||
|
||||
### 方式 A:用配套脚本(推荐,一步到位)
|
||||
|
||||
```bash
|
||||
# 扫描仓库 Issue,输出新手任务看板(Markdown)
|
||||
python scripts/newcomer.py --owner Gitlink --repo gitlink-cli
|
||||
|
||||
# 输出 JSON 供 Agent 进一步处理(含每个候选的引导文案)
|
||||
python scripts/newcomer.py --owner Gitlink --repo gitlink-cli --format json
|
||||
|
||||
# 写入文件
|
||||
python scripts/newcomer.py --owner Gitlink --repo gitlink-cli --output board.md
|
||||
```
|
||||
|
||||
参数说明:
|
||||
|
||||
| 参数 | 类型 | 必填 | 说明 |
|
||||
|------|------|:----:|------|
|
||||
| `--owner` | string | 是* | 仓库所有者(*或用 `--slug`) |
|
||||
| `--repo` | string | 是* | 仓库名称 |
|
||||
| `--slug` | string | 否 | `owner/repo` 或完整 URL,替代 owner/repo |
|
||||
| `--issue` | int | 否 | 只为指定 Issue 编号(web 序号)生成引导 |
|
||||
| `--limit` | int | 否 | 扫描 Issue 数量上限,默认 50 |
|
||||
| `--format` | string | 否 | `markdown`(默认)或 `json` |
|
||||
| `--output` | string | 否 | 输出文件路径,缺省打印到标准输出 |
|
||||
|
||||
### 方式 B:用 gitlink-cli 命令手动采集
|
||||
|
||||
当无法运行脚本时,Agent 可用以下命令采集数据后自行分析:
|
||||
|
||||
```bash
|
||||
# 1. 获取开放 Issue 列表
|
||||
gitlink-cli issue +list --owner Gitlink --repo gitlink-cli --state open --format json
|
||||
|
||||
# 2. 查看某个 Issue 详情(--number 用 web 序号,不是全局 id)
|
||||
gitlink-cli issue +view --owner Gitlink --repo gitlink-cli --number 12 --format json
|
||||
|
||||
# 3. 查看仓库标签,确认是否有 good first issue 类标签
|
||||
gitlink-cli label +list --owner Gitlink --repo gitlink-cli --format json
|
||||
```
|
||||
|
||||
识别规则建议:
|
||||
- 带 `good first issue` / `beginner` / `新手` 等标签 → 强新手信号
|
||||
- 标题/正文含 `typo` / `docs` / `文档` / `test` / `翻译` 等 → 易上手信号
|
||||
- 含 `refactor` / `架构` / `并发` / `性能` 等 → 高难度信号
|
||||
|
||||
## 工作流 2:为单个 Issue 生成引导评论
|
||||
|
||||
```bash
|
||||
# 生成引导文案(不发布)
|
||||
python scripts/newcomer.py --owner Gitlink --repo gitlink-cli --issue 12
|
||||
|
||||
# 确认文案后,由用户决定是否发布为评论(写操作,需确认)
|
||||
gitlink-cli issue +comment --owner Gitlink --repo gitlink-cli --number 12 -b "<引导文案>"
|
||||
```
|
||||
|
||||
## API 注意事项
|
||||
|
||||
- **ID 混淆**:GitLink 的 Issue 列表接口(`issue +list`)通常只返回全局数据库 `id`,不含 web 序号。发评论 / 关联 Issue 时,`issue +comment` 的 `--number` 需要 web 序号(URL 中显示的编号),不要把全局 id 当 web 序号用。
|
||||
- **写操作确认**:`issue +comment` 会真实发布评论,执行前务必向用户确认内容与目标 Issue。
|
||||
- 数据采集全程只读,脚本默认不发布任何内容。
|
||||
|
||||
## 输出示例
|
||||
|
||||
参见 [`examples/`](examples/) 目录下的真实运行产物(新手看板、单 Issue 引导)。
|
||||
|
||||
## References
|
||||
|
||||
- [api-reference.md](references/api-reference.md) — 采集的接口、字段与 ID 混淆说明
|
||||
- [scoring.md](references/scoring.md) — 新手友好度评分规则与信号词表
|
||||
- [gitlink-shared](../gitlink-shared/SKILL.md) — 认证、全局参数、安全规则
|
||||
|
|
@ -0,0 +1,64 @@
|
|||
# gitlink-newcomer API 参考
|
||||
|
||||
> **前置条件:** 先阅读 [`../../gitlink-shared/SKILL.md`](../../gitlink-shared/SKILL.md) 了解认证、全局参数和安全规则。
|
||||
|
||||
本技能识别 good-first-issue 所依赖的 GitLink 接口与字段。数据采集全程只读。
|
||||
|
||||
## 采集的接口
|
||||
|
||||
### Issue 列表
|
||||
|
||||
```
|
||||
GET /:owner/:repo/issues.json?page={page}&limit={limit}
|
||||
# 或经 gitlink-cli:
|
||||
gitlink-cli issue +list --owner <owner> --repo <repo> --state open --format json
|
||||
```
|
||||
|
||||
返回的 `issues[]` 中本技能使用的字段:
|
||||
|
||||
| 字段 | 说明 | 用途 |
|
||||
|------|------|------|
|
||||
| `id` | 全局数据库 ID | 去重、展示(**非** web 序号) |
|
||||
| `name` / `subject` | Issue 标题 | 关键词分析、展示 |
|
||||
| `description` | 正文 | 关键词分析、长度评估 |
|
||||
| `issue_tags` / `labels` | 标签 | good-first-issue 标签信号 |
|
||||
| `comment_journals_count` / `journals_count` | 评论数 | 讨论热度评估 |
|
||||
| `author_login` / `author_name` | 作者 | 展示 |
|
||||
|
||||
### Issue 详情(单个)
|
||||
|
||||
```
|
||||
GET /:owner/:repo/issues/{number}.json
|
||||
# 或:
|
||||
gitlink-cli issue +view --owner <owner> --repo <repo> --number {number} --format json
|
||||
```
|
||||
|
||||
`{number}` 为 web 序号(URL 中的编号)。详情接口可拿到更完整的字段。
|
||||
|
||||
## ID 混淆(重要)
|
||||
|
||||
GitLink 存在两种 ID:
|
||||
|
||||
| 名称 | 来源 | 用途 |
|
||||
|------|------|------|
|
||||
| 全局数据库 `id` | Issue 列表接口 | 仅去重 / 内部引用 |
|
||||
| web 序号(`project_issues_index` / `number`) | 单 Issue 详情、web URL | 发评论、PR 关联、`issue +comment --number` |
|
||||
|
||||
本技能的处理:
|
||||
- 列表接口只返回全局 `id` 时,看板以 `id:xxx` 形式标注,**不**伪装成 web 序号;
|
||||
- 生成引导评论时,若无可靠 web 序号则用 Issue 标题引用,避免误导;
|
||||
- 发布评论用 `gitlink-cli issue +comment --number <web序号>`。
|
||||
|
||||
## 写操作
|
||||
|
||||
发布引导评论是写操作:
|
||||
|
||||
```
|
||||
gitlink-cli issue +comment --owner <owner> --repo <repo> --number <web序号> -b "<引导文案>"
|
||||
```
|
||||
|
||||
执行前必须征得用户确认;本技能默认只生成文案,不自动发布。
|
||||
|
||||
## 错误处理
|
||||
|
||||
沿用 gitlink-shared 的错误码(401 重新登录 / 403 权限 / 404 检查 owner/repo)。采集失败时脚本返回非零退出码并打印原因。
|
||||
|
|
@ -0,0 +1,42 @@
|
|||
# 新手友好度评分规则
|
||||
|
||||
本技能用一套可解释的规则为每个 Issue 打出 0-100 的新手友好度分,不依赖大模型。
|
||||
|
||||
## 评分模型
|
||||
|
||||
基准分 **50**,在此基础上根据信号加减:
|
||||
|
||||
| 信号 | 分值 | 说明 |
|
||||
|------|:----:|------|
|
||||
| 命中新手友好标签 | +35 | `good first issue` / `beginner` / `新手` / `easy` 等 |
|
||||
| 命中高难度标签 | -30 | `hard` / `complex` / `advanced` / `困难` 等 |
|
||||
| 命中易上手关键词 | +7/词(上限 +20) | `typo` / `docs` / `文档` / `test` / `翻译` 等 |
|
||||
| 命中高难度关键词 | -10/词(上限 -25) | `refactor` / `架构` / `并发` / `性能` / `安全` 等 |
|
||||
| 描述长度适中(30-600) | +5 | 适中描述更易上手 |
|
||||
| 描述过长(>1500) | -8 | 往往较复杂 |
|
||||
| 讨论过多(评论 >15) | -8 | 可能存在分歧 |
|
||||
|
||||
最终分数裁剪到 [0, 100]。
|
||||
|
||||
## 难度等级
|
||||
|
||||
| 友好度分 | 难度等级 |
|
||||
|:--------:|:--------:|
|
||||
| ≥ 75 | 入门 |
|
||||
| 55 - 74 | 较易 |
|
||||
| 40 - 54 | 中等 |
|
||||
| < 40 | 进阶 |
|
||||
|
||||
友好度 ≥ 55 的 Issue 被列为新手候选(`is_good_first = true`)。
|
||||
|
||||
## 信号词表(节选)
|
||||
|
||||
**新手友好标签**:good first issue、good-first-issue、first-timers-only、beginner、beginner-friendly、easy、starter、新手、新手友好、入门、简单
|
||||
|
||||
**易上手关键词**:typo、document、docs、readme、comment、translation、rename、format、lint、test、example、i18n、文档、注释、拼写、翻译、示例、格式
|
||||
|
||||
**高难度关键词**:refactor、architecture、performance、concurrency、race、security、deadlock、memory leak、breaking change、重构、架构、性能、并发、安全、死锁、内存
|
||||
|
||||
## 可调整性
|
||||
|
||||
词表与阈值集中在 `scripts/newcomer.py` 顶部常量(`GOOD_FIRST_LABELS` / `EASY_KEYWORDS` / `HARD_KEYWORDS` / `HARD_LABELS`),可按项目习惯调整。
|
||||
|
|
@ -0,0 +1,241 @@
|
|||
"""GitLink 公开 API 共享客户端。
|
||||
|
||||
供 gitlink-skills-pack 下各 Skill 的脚本复用。仅依赖 Python 标准库,
|
||||
无需第三方包,便于在受限环境或 Agent 沙箱中运行。
|
||||
|
||||
数据全部来自 GitLink 平台公开接口(https://www.gitlink.org.cn/api),
|
||||
默认无需 token;如需访问私有仓库,可传入 token。
|
||||
|
||||
所有方法均为只读,不修改任何远程数据。
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import base64
|
||||
import json
|
||||
import time
|
||||
import urllib.error
|
||||
import urllib.parse
|
||||
import urllib.request
|
||||
from pathlib import Path
|
||||
from typing import Any
|
||||
|
||||
API_BASE = "https://www.gitlink.org.cn/api"
|
||||
USER_AGENT = "gitlink-skills-pack/1.0 (+https://www.gitlink.org.cn)"
|
||||
DEFAULT_TIMEOUT = 30
|
||||
COMMIT_PAGE_SIZE = 50 # GitLink commits 接口每页硬上限
|
||||
|
||||
|
||||
class GitLinkError(RuntimeError):
|
||||
"""API 调用中不可恢复的错误。"""
|
||||
|
||||
|
||||
class GitLinkClient:
|
||||
"""GitLink 公开数据接口客户端。
|
||||
|
||||
带可选文件缓存:同一资源重复读取不重复打网,对平台友好。
|
||||
"""
|
||||
|
||||
def __init__(self, base: str = API_BASE, token: str | None = None,
|
||||
timeout: int = DEFAULT_TIMEOUT, cache_dir: Path | None = None) -> None:
|
||||
self.base = base.rstrip("/")
|
||||
self.token = token
|
||||
self.timeout = timeout
|
||||
self.cache_dir = cache_dir
|
||||
if self.cache_dir:
|
||||
self.cache_dir.mkdir(parents=True, exist_ok=True)
|
||||
|
||||
# ------------------------------------------------------------------
|
||||
# 底层请求
|
||||
# ------------------------------------------------------------------
|
||||
def _cache_path(self, url: str) -> Path | None:
|
||||
if not self.cache_dir:
|
||||
return None
|
||||
safe = urllib.parse.quote(url, safe="")
|
||||
return self.cache_dir / f"{safe}.json"
|
||||
|
||||
def get(self, path: str, query: dict[str, Any] | None = None) -> Any:
|
||||
"""GET 请求,返回解析后的 JSON(dict/list)或 None。"""
|
||||
url = f"{self.base}/{path.lstrip('/')}"
|
||||
if query:
|
||||
url = f"{url}?{urllib.parse.urlencode(query)}"
|
||||
|
||||
cache_path = self._cache_path(url)
|
||||
if cache_path and cache_path.exists():
|
||||
return json.loads(cache_path.read_text(encoding="utf-8"))
|
||||
|
||||
headers = {"Accept": "application/json", "User-Agent": USER_AGENT}
|
||||
if self.token:
|
||||
headers["Authorization"] = f"Bearer {self.token}"
|
||||
|
||||
req = urllib.request.Request(url, headers=headers)
|
||||
try:
|
||||
with urllib.request.urlopen(req, timeout=self.timeout) as resp:
|
||||
raw = resp.read().decode("utf-8", errors="replace")
|
||||
except urllib.error.HTTPError as exc:
|
||||
raise GitLinkError(f"HTTP {exc.code}: {url}") from exc
|
||||
except urllib.error.URLError as exc:
|
||||
raise GitLinkError(f"网络错误: {url} -> {exc.reason}") from exc
|
||||
|
||||
text = raw.strip()
|
||||
if not text or text in ("null", "{}", "[]"):
|
||||
data: Any = None
|
||||
elif text[0] in "{[":
|
||||
try:
|
||||
data = json.loads(text)
|
||||
except json.JSONDecodeError as exc:
|
||||
raise GitLinkError(f"响应非 JSON: {url}") from exc
|
||||
else:
|
||||
raise GitLinkError(f"响应非 JSON(可能是 HTML): {url}")
|
||||
|
||||
if cache_path is not None:
|
||||
cache_path.write_text(json.dumps(data, ensure_ascii=False), encoding="utf-8")
|
||||
return data
|
||||
|
||||
# ------------------------------------------------------------------
|
||||
# 资源访问(高层封装)
|
||||
# ------------------------------------------------------------------
|
||||
def repo_info(self, owner: str, repo: str) -> dict[str, Any]:
|
||||
"""仓库元信息。"""
|
||||
data = self.get(f"{owner}/{repo}.json")
|
||||
return data if isinstance(data, dict) else {}
|
||||
|
||||
def issues(self, owner: str, repo: str, limit: int = 50,
|
||||
page: int = 1) -> list[dict[str, Any]]:
|
||||
"""Issue 列表。"""
|
||||
data = self.get(f"{owner}/{repo}/issues.json", {"page": page, "limit": limit})
|
||||
return _extract_list(data, ("issues",))
|
||||
|
||||
def issue_detail(self, owner: str, repo: str, number: int) -> dict[str, Any]:
|
||||
"""单个 Issue 详情(含完整字段)。"""
|
||||
data = self.get(f"{owner}/{repo}/issues/{number}.json")
|
||||
return data if isinstance(data, dict) else {}
|
||||
|
||||
def pulls(self, owner: str, repo: str, limit: int = 50,
|
||||
page: int = 1) -> list[dict[str, Any]]:
|
||||
"""PR 列表。"""
|
||||
data = self.get(f"{owner}/{repo}/pulls.json", {"page": page, "limit": limit})
|
||||
return _extract_list(data, ("issues", "pulls"))
|
||||
|
||||
def contributors(self, owner: str, repo: str) -> list[dict[str, Any]]:
|
||||
"""贡献者列表。"""
|
||||
data = self.get(f"{owner}/{repo}/contributors.json")
|
||||
return _extract_list(data, ("list",))
|
||||
|
||||
def commits(self, owner: str, repo: str, max_pages: int = 4) -> list[dict[str, Any]]:
|
||||
"""提交列表(按需翻页,每页 50 条,以 total_count 为终止依据)。"""
|
||||
out: list[dict[str, Any]] = []
|
||||
total: int | None = None
|
||||
for page in range(1, max(1, max_pages) + 1):
|
||||
data = self.get(f"{owner}/{repo}/commits.json",
|
||||
{"page": page, "limit": COMMIT_PAGE_SIZE})
|
||||
if total is None and isinstance(data, dict):
|
||||
total = _safe_int(data.get("total_count")) or None
|
||||
page_items = _extract_list(data, ("commits",))
|
||||
if not page_items:
|
||||
break
|
||||
out.extend(page_items)
|
||||
if total is not None and len(out) >= total:
|
||||
break
|
||||
return out
|
||||
|
||||
def list_dir(self, owner: str, repo: str, path: str = "",
|
||||
ref: str = "master") -> list[dict[str, Any]]:
|
||||
"""列出目录下的条目(文件与子目录)。
|
||||
|
||||
返回的每个 entry 含 name / path / type(file|dir) / sha / size,
|
||||
文件类型的 entry 还可能直接带明文 content。
|
||||
"""
|
||||
data = self.get(f"{owner}/{repo}/sub_entries.json",
|
||||
{"filepath": path, "ref": ref})
|
||||
# 查询目录时 entries 为 list;查询单文件时 entries 为单个 dict。
|
||||
# 统一归一化为 list,便于下游处理。
|
||||
if isinstance(data, dict):
|
||||
entries = data.get("entries")
|
||||
if isinstance(entries, dict):
|
||||
return [entries]
|
||||
if isinstance(entries, list):
|
||||
return entries
|
||||
return _extract_list(data, ("entries",))
|
||||
|
||||
def file_content(self, owner: str, repo: str, filepath: str,
|
||||
ref: str = "master") -> str | None:
|
||||
"""读取单个文件的文本内容。
|
||||
|
||||
GitLink 的 sub_entries 接口对单文件查询会在 entries 中返回明文 content,
|
||||
据此取出。文件不存在或无内容时返回 None。
|
||||
"""
|
||||
entries = self.list_dir(owner, repo, filepath, ref)
|
||||
target = filepath.rsplit("/", 1)[-1]
|
||||
for entry in entries:
|
||||
if entry.get("type") == "file" and entry.get("name") == target:
|
||||
content = entry.get("content")
|
||||
if isinstance(content, str):
|
||||
return content
|
||||
# 回退:部分情况下单文件查询 entries 仅一项
|
||||
if len(entries) == 1 and entries[0].get("type") == "file":
|
||||
content = entries[0].get("content")
|
||||
if isinstance(content, str):
|
||||
return content
|
||||
return None
|
||||
|
||||
def readme(self, owner: str, repo: str, ref: str = "master") -> str | None:
|
||||
"""读取仓库 README(自动 base64 解码)。"""
|
||||
data = self.get(f"{owner}/{repo}/readme.json", {"ref": ref})
|
||||
if not isinstance(data, dict):
|
||||
return None
|
||||
content = data.get("content")
|
||||
if not isinstance(content, str):
|
||||
return None
|
||||
# 注意:GitLink 的 readme.json 虽然 encoding 标为 base64,
|
||||
# 实测 content 多为明文 Markdown。先探测明文特征,命中则直接返回;
|
||||
# 否则再尝试 base64 解码。
|
||||
stripped = content.lstrip()
|
||||
if stripped.startswith(("#", "<", "[", "-", "*", "本", "这", "项")) or "\n" in content[:200]:
|
||||
return content
|
||||
try:
|
||||
raw = base64.b64decode(content.encode("ascii", "ignore"))
|
||||
decoded = raw.decode("utf-8", errors="replace")
|
||||
# 解码结果若不像文本(大量替换符),回退为原文
|
||||
if decoded.count("\ufffd") > len(decoded) * 0.1:
|
||||
return content
|
||||
return decoded
|
||||
except (ValueError, TypeError):
|
||||
return content
|
||||
|
||||
|
||||
# ----------------------------------------------------------------------------
|
||||
# 辅助
|
||||
# ----------------------------------------------------------------------------
|
||||
|
||||
def _extract_list(payload: Any, keys: tuple[str, ...]) -> list[Any]:
|
||||
"""从可能嵌套的响应中提取第一个匹配键的列表。"""
|
||||
if isinstance(payload, list):
|
||||
return payload
|
||||
if isinstance(payload, dict):
|
||||
for key in keys:
|
||||
value = payload.get(key)
|
||||
if isinstance(value, list):
|
||||
return value
|
||||
return []
|
||||
|
||||
|
||||
def _safe_int(value: Any, default: int = 0) -> int:
|
||||
try:
|
||||
return int(value)
|
||||
except (TypeError, ValueError):
|
||||
return default
|
||||
|
||||
|
||||
def split_owner_repo(slug: str) -> tuple[str, str]:
|
||||
"""把 'owner/repo' 或完整 URL 解析为 (owner, repo)。"""
|
||||
s = slug.strip()
|
||||
if s.startswith("http"):
|
||||
parts = urllib.parse.urlparse(s).path.strip("/").split("/")
|
||||
if len(parts) >= 2:
|
||||
return parts[0], parts[1].replace(".git", "")
|
||||
raise GitLinkError(f"无法从 URL 解析 owner/repo: {slug}")
|
||||
if "/" in s:
|
||||
owner, repo = s.split("/", 1)
|
||||
return owner, repo.replace(".git", "")
|
||||
raise GitLinkError(f"格式应为 owner/repo: {slug}")
|
||||
|
|
@ -0,0 +1,333 @@
|
|||
"""gitlink-newcomer:新人引导分析。
|
||||
|
||||
识别一个 GitLink 仓库中适合新贡献者上手的 Issue,评估上手难度与友好度,
|
||||
为每个候选 Issue 生成个性化引导评论,并产出新手任务看板。
|
||||
|
||||
数据来自 GitLink 公开 API(只读),无需登录。生成的引导评论仅作为建议输出,
|
||||
是否发布由用户通过 gitlink-cli 自行决定。
|
||||
|
||||
用法:
|
||||
python newcomer.py --owner Gitlink --repo gitlink-cli
|
||||
python newcomer.py --owner Gitlink --repo gitlink-cli --format json
|
||||
python newcomer.py --owner Gitlink --repo gitlink-cli --issue 12 # 只为某个 Issue 生成引导
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import argparse
|
||||
import json
|
||||
import re
|
||||
import sys
|
||||
from pathlib import Path
|
||||
from typing import Any
|
||||
|
||||
sys.path.insert(0, str(Path(__file__).resolve().parent))
|
||||
from glapi import GitLinkClient, GitLinkError, split_owner_repo
|
||||
|
||||
# Windows 控制台默认 GBK,直接打印含 emoji 的 Markdown 会抛 UnicodeEncodeError。
|
||||
# 重配置 stdout 为 UTF-8,确保跨平台正常输出。
|
||||
if hasattr(sys.stdout, "reconfigure"):
|
||||
try:
|
||||
sys.stdout.reconfigure(encoding="utf-8")
|
||||
except Exception:
|
||||
pass
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# 识别规则
|
||||
# ---------------------------------------------------------------------------
|
||||
|
||||
# good-first-issue 的标签信号(不同项目命名习惯)
|
||||
GOOD_FIRST_LABELS = {
|
||||
"good first issue", "good-first-issue", "goodfirstissue",
|
||||
"good first", "first-timers-only", "first timers only",
|
||||
"beginner", "beginner-friendly", "easy", "starter",
|
||||
"新手", "新手友好", "新人", "入门", "简单",
|
||||
}
|
||||
|
||||
# 标题/正文中暗示「适合新手」的关键词
|
||||
EASY_KEYWORDS = [
|
||||
"typo", "document", "docs", "readme", "comment", "translation", "translate",
|
||||
"rename", "format", "lint", "test", "example", "i18n",
|
||||
"文档", "注释", "拼写", "翻译", "示例", "格式", "重命名",
|
||||
]
|
||||
|
||||
# 暗示「难度较高、不适合新手」的关键词
|
||||
HARD_KEYWORDS = [
|
||||
"refactor", "architecture", "performance", "concurrency", "race",
|
||||
"security", "deadlock", "memory leak", "breaking change",
|
||||
"重构", "架构", "性能", "并发", "安全", "死锁", "内存",
|
||||
]
|
||||
|
||||
# 难度高的标签
|
||||
HARD_LABELS = {"hard", "complex", "advanced", "epic", "困难", "复杂"}
|
||||
|
||||
|
||||
def _text_of(issue: dict[str, Any]) -> str:
|
||||
"""合并 Issue 的标题与正文用于关键词分析。"""
|
||||
title = issue.get("name") or issue.get("subject") or issue.get("title") or ""
|
||||
body = issue.get("description") or issue.get("body") or ""
|
||||
return f"{title}\n{body}".lower()
|
||||
|
||||
|
||||
def _labels_of(issue: dict[str, Any]) -> list[str]:
|
||||
"""提取 Issue 标签名(兼容多种字段结构)。"""
|
||||
labels: list[str] = []
|
||||
raw = issue.get("issue_tags") or issue.get("labels") or issue.get("tags")
|
||||
if isinstance(raw, list):
|
||||
for item in raw:
|
||||
if isinstance(item, dict):
|
||||
name = item.get("name") or item.get("title")
|
||||
if name:
|
||||
labels.append(str(name).lower())
|
||||
elif item:
|
||||
labels.append(str(item).lower())
|
||||
return labels
|
||||
|
||||
|
||||
def _issue_number(issue: dict[str, Any]) -> int | None:
|
||||
"""提取仓库内 Issue 序号(web URL 中显示的编号)。
|
||||
|
||||
注意:GitLink 的 Issue 列表接口通常只返回全局数据库 id,不含 web 序号;
|
||||
只有单 Issue 详情或带 number/index 字段时才有可靠序号。取不到返回 None,
|
||||
避免把全局 id 误当作 web 序号引用。
|
||||
"""
|
||||
for key in ("number", "index"):
|
||||
val = issue.get(key)
|
||||
if isinstance(val, int):
|
||||
return val
|
||||
return None
|
||||
|
||||
|
||||
def _issue_gid(issue: dict[str, Any]) -> int | None:
|
||||
"""全局数据库 id(仅用于去重/展示,不可用于 web 链接 / PR 关联)。"""
|
||||
val = issue.get("id")
|
||||
return val if isinstance(val, int) else None
|
||||
|
||||
|
||||
def score_issue(issue: dict[str, Any]) -> dict[str, Any]:
|
||||
"""评估单个 Issue 的新手友好度。
|
||||
|
||||
返回友好度评分(0-100)、难度等级、命中的信号,
|
||||
评分越高越适合新人上手。
|
||||
"""
|
||||
text = _text_of(issue)
|
||||
labels = _labels_of(issue)
|
||||
signals: list[str] = []
|
||||
score = 50 # 基准分
|
||||
|
||||
# 标签信号(最强)
|
||||
if any(lb in GOOD_FIRST_LABELS for lb in labels):
|
||||
score += 35
|
||||
signals.append("带有新手友好标签")
|
||||
if any(lb in HARD_LABELS for lb in labels):
|
||||
score -= 30
|
||||
signals.append("带有高难度标签")
|
||||
|
||||
# 关键词信号
|
||||
easy_hits = [k for k in EASY_KEYWORDS if k in text]
|
||||
if easy_hits:
|
||||
score += min(20, len(easy_hits) * 7)
|
||||
signals.append(f"内容涉及易上手主题({', '.join(easy_hits[:3])})")
|
||||
hard_hits = [k for k in HARD_KEYWORDS if k in text]
|
||||
if hard_hits:
|
||||
score -= min(25, len(hard_hits) * 10)
|
||||
signals.append(f"内容涉及高难度主题({', '.join(hard_hits[:3])})")
|
||||
|
||||
# 描述长度:太长往往复杂
|
||||
body = issue.get("description") or issue.get("body") or ""
|
||||
if len(body) > 1500:
|
||||
score -= 8
|
||||
signals.append("描述较长,可能较复杂")
|
||||
elif 30 <= len(body) <= 600:
|
||||
score += 5
|
||||
signals.append("描述长度适中")
|
||||
|
||||
# 评论数:讨论太多可能有争议或难度大
|
||||
comments = issue.get("comment_journals_count") or issue.get("journals_count") or 0
|
||||
if isinstance(comments, int) and comments > 15:
|
||||
score -= 8
|
||||
signals.append("讨论较多,可能存在分歧")
|
||||
|
||||
score = max(0, min(100, score))
|
||||
if score >= 75:
|
||||
difficulty = "入门"
|
||||
elif score >= 55:
|
||||
difficulty = "较易"
|
||||
elif score >= 40:
|
||||
difficulty = "中等"
|
||||
else:
|
||||
difficulty = "进阶"
|
||||
|
||||
return {
|
||||
"number": _issue_number(issue),
|
||||
"gid": _issue_gid(issue),
|
||||
"title": issue.get("name") or issue.get("subject") or issue.get("title") or "(无标题)",
|
||||
"labels": labels,
|
||||
"friendliness": score,
|
||||
"difficulty": difficulty,
|
||||
"signals": signals,
|
||||
"author": issue.get("author_login") or issue.get("author_name"),
|
||||
"comments": comments if isinstance(comments, int) else 0,
|
||||
"is_good_first": score >= 55,
|
||||
}
|
||||
|
||||
|
||||
def build_guidance(scored: dict[str, Any], owner: str, repo: str) -> str:
|
||||
"""为一个候选 Issue 生成个性化的新手引导评论。"""
|
||||
num = scored.get("number")
|
||||
title = scored["title"]
|
||||
diff = scored["difficulty"]
|
||||
# 有可靠 web 序号时用 #num 引用,否则用标题引用(不误用全局 id)
|
||||
ref = f"#{num}" if num else f"《{title}》"
|
||||
|
||||
lines = [
|
||||
f"👋 欢迎!这个 Issue({ref})被 gitlink-newcomer 评估为 **{diff}** 难度,适合作为参与本项目的起点。",
|
||||
"",
|
||||
"如果你想认领它,建议按以下步骤上手:",
|
||||
"",
|
||||
f"1. 阅读项目的 `README` 和 `CONTRIBUTING`(如有),了解开发与提交规范。",
|
||||
f"2. Fork 本仓库并克隆你的 Fork:`gitlink-cli repo +fork --owner {owner} --repo {repo}`",
|
||||
"3. 新建一个分支进行修改,保持改动聚焦于本 Issue。",
|
||||
f"4. 完成后从你的 Fork 向 `{owner}/{repo}` 提交 PR,并在描述里关联本 Issue。",
|
||||
"",
|
||||
]
|
||||
|
||||
if "文档" in str(scored["signals"]) or "docs" in str(scored["signals"]).lower():
|
||||
lines.append("> 提示:这看起来是一个文档/文本类改动,通常不需要改动核心逻辑,很适合第一次贡献。")
|
||||
if scored["comments"] > 8:
|
||||
lines.append("> 提示:该 Issue 已有较多讨论,动手前建议先通读评论,确认当前结论与分工。")
|
||||
lines.append("")
|
||||
lines.append("有任何问题都可以在本 Issue 下留言,社区很乐意帮助新人。祝贡献顺利!🚀")
|
||||
return "\n".join(lines)
|
||||
|
||||
|
||||
def analyze(owner: str, repo: str, limit: int = 50,
|
||||
client: GitLinkClient | None = None) -> dict[str, Any]:
|
||||
"""分析仓库的 Issue,识别并排序新手友好的候选。"""
|
||||
client = client or GitLinkClient()
|
||||
issues = client.issues(owner, repo, limit=limit)
|
||||
scored = [score_issue(i) for i in issues]
|
||||
candidates = [s for s in scored if s["is_good_first"]]
|
||||
candidates.sort(key=lambda s: s["friendliness"], reverse=True)
|
||||
|
||||
return {
|
||||
"owner": owner,
|
||||
"repo": repo,
|
||||
"total_issues": len(issues),
|
||||
"candidate_count": len(candidates),
|
||||
"candidates": candidates,
|
||||
"all_scored": scored,
|
||||
}
|
||||
|
||||
|
||||
def render_board(result: dict[str, Any], owner: str, repo: str) -> str:
|
||||
"""渲染新手任务看板(Markdown)。"""
|
||||
lines = [
|
||||
f"# 新手任务看板 — {owner}/{repo}",
|
||||
"",
|
||||
f"由 gitlink-newcomer 生成。共扫描 {result['total_issues']} 个开放 Issue,"
|
||||
f"识别出 **{result['candidate_count']}** 个适合新贡献者上手的任务。",
|
||||
"",
|
||||
]
|
||||
if not result["candidates"]:
|
||||
lines += [
|
||||
"暂未发现明显适合新手的 Issue。建议维护者:",
|
||||
"",
|
||||
"- 为简单任务打上 `good first issue` 标签",
|
||||
"- 在 Issue 描述里补充清晰的上手说明与验收标准",
|
||||
"",
|
||||
]
|
||||
return "\n".join(lines)
|
||||
|
||||
lines += [
|
||||
"| 推荐度 | 难度 | Issue | 标题 | 命中信号 |",
|
||||
"|:------:|:----:|:-----:|------|----------|",
|
||||
]
|
||||
for c in result["candidates"]:
|
||||
stars = "⭐" * max(1, round(c["friendliness"] / 20))
|
||||
signal = c["signals"][0] if c["signals"] else "-"
|
||||
title = c["title"][:40]
|
||||
# 有 web 序号用 #num,否则标注全局 id(gid)以便定位
|
||||
if c.get("number"):
|
||||
ref = f"#{c['number']}"
|
||||
elif c.get("gid"):
|
||||
ref = f"id:{c['gid']}"
|
||||
else:
|
||||
ref = "-"
|
||||
lines.append(
|
||||
f"| {stars} | {c['difficulty']} | {ref} | {title} | {signal} |"
|
||||
)
|
||||
lines += [
|
||||
"",
|
||||
"## 建议行动",
|
||||
"",
|
||||
"1. 为上述 Issue 添加引导评论,欢迎新贡献者认领(见各 Issue 的引导文案)。",
|
||||
"2. 确认这些 Issue 的描述包含足够的上手信息。",
|
||||
"3. 可在仓库 README 中链接本看板,方便新人发现。",
|
||||
"",
|
||||
]
|
||||
return "\n".join(lines)
|
||||
|
||||
|
||||
def main(argv: list[str] | None = None) -> int:
|
||||
p = argparse.ArgumentParser(
|
||||
prog="gitlink-newcomer",
|
||||
description="识别 good-first-issue 并生成新手引导",
|
||||
)
|
||||
p.add_argument("--owner", help="仓库所有者,如 Gitlink")
|
||||
p.add_argument("--repo", help="仓库名称,如 gitlink-cli")
|
||||
p.add_argument("--slug", help="owner/repo 形式,或完整仓库 URL")
|
||||
p.add_argument("--issue", type=int, help="只为指定 Issue 编号生成引导评论")
|
||||
p.add_argument("--limit", type=int, default=50, help="扫描的 Issue 数量上限(默认 50)")
|
||||
p.add_argument("--format", choices=["markdown", "json"], default="markdown",
|
||||
help="输出格式(默认 markdown)")
|
||||
p.add_argument("--output", type=Path, help="输出文件路径,缺省打印到标准输出")
|
||||
args = p.parse_args(argv)
|
||||
|
||||
if args.slug:
|
||||
owner, repo = split_owner_repo(args.slug)
|
||||
elif args.owner and args.repo:
|
||||
owner, repo = args.owner, args.repo
|
||||
else:
|
||||
print("错误:请用 --owner/--repo 或 --slug 指定仓库。", file=sys.stderr)
|
||||
return 2
|
||||
|
||||
client = GitLinkClient()
|
||||
try:
|
||||
# 单 Issue 引导模式
|
||||
if args.issue is not None:
|
||||
detail = client.issue_detail(owner, repo, args.issue)
|
||||
if not detail:
|
||||
print(f"未找到 Issue #{args.issue}", file=sys.stderr)
|
||||
return 1
|
||||
scored = score_issue(detail)
|
||||
guidance = build_guidance(scored, owner, repo)
|
||||
if args.format == "json":
|
||||
out = json.dumps({"issue": scored, "guidance": guidance},
|
||||
ensure_ascii=False, indent=2)
|
||||
else:
|
||||
out = guidance
|
||||
else:
|
||||
result = analyze(owner, repo, limit=args.limit, client=client)
|
||||
if args.format == "json":
|
||||
# 为每个候选附上引导文案
|
||||
for c in result["candidates"]:
|
||||
c["guidance"] = build_guidance(c, owner, repo)
|
||||
out = json.dumps(result, ensure_ascii=False, indent=2)
|
||||
else:
|
||||
out = render_board(result, owner, repo)
|
||||
except GitLinkError as exc:
|
||||
print(f"采集失败:{exc}", file=sys.stderr)
|
||||
return 1
|
||||
|
||||
if args.output:
|
||||
args.output.parent.mkdir(parents=True, exist_ok=True)
|
||||
args.output.write_text(out, encoding="utf-8")
|
||||
print(f"已写入 {args.output}")
|
||||
else:
|
||||
print(out)
|
||||
return 0
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
raise SystemExit(main())
|
||||
|
|
@ -0,0 +1,128 @@
|
|||
"""gitlink-newcomer 单元测试。
|
||||
|
||||
覆盖 good-first-issue 识别、友好度评分、ID 区分、引导生成与看板渲染。
|
||||
使用合成数据,不触网,可离线运行:
|
||||
|
||||
python -m pytest tests/ -v
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import sys
|
||||
from pathlib import Path
|
||||
|
||||
sys.path.insert(0, str(Path(__file__).resolve().parent.parent / "scripts"))
|
||||
|
||||
import pytest
|
||||
|
||||
from newcomer import (
|
||||
score_issue, build_guidance, render_board,
|
||||
_issue_number, _issue_gid, _labels_of,
|
||||
)
|
||||
|
||||
|
||||
def issue(name="", body="", labels=None, comments=0, number=None, gid=None):
|
||||
d = {"name": name, "description": body, "comment_journals_count": comments}
|
||||
if labels is not None:
|
||||
d["issue_tags"] = [{"name": x} for x in labels]
|
||||
if number is not None:
|
||||
d["number"] = number
|
||||
if gid is not None:
|
||||
d["id"] = gid
|
||||
return d
|
||||
|
||||
|
||||
class TestLabels:
|
||||
def test_extract_dict_labels(self):
|
||||
assert _labels_of(issue(labels=["Bug", "good first issue"])) == ["bug", "good first issue"]
|
||||
|
||||
def test_no_labels(self):
|
||||
assert _labels_of(issue()) == []
|
||||
|
||||
|
||||
class TestIdDistinction:
|
||||
def test_web_number_preferred(self):
|
||||
assert _issue_number(issue(number=12)) == 12
|
||||
|
||||
def test_no_web_number_returns_none(self):
|
||||
# 列表接口只有全局 id,不应被当作 web 序号
|
||||
assert _issue_number(issue(gid=140801)) is None
|
||||
|
||||
def test_gid_extracted(self):
|
||||
assert _issue_gid(issue(gid=140801)) == 140801
|
||||
|
||||
|
||||
class TestScoreIssue:
|
||||
def test_good_first_label_boosts(self):
|
||||
s = score_issue(issue(name="fix typo", labels=["good first issue"]))
|
||||
assert s["friendliness"] >= 75
|
||||
assert s["is_good_first"] is True
|
||||
assert "新手友好标签" in str(s["signals"])
|
||||
|
||||
def test_hard_keyword_lowers(self):
|
||||
s = score_issue(issue(name="refactor concurrency architecture",
|
||||
body="needs deep refactor of the core"))
|
||||
assert s["friendliness"] < 55
|
||||
assert s["is_good_first"] is False
|
||||
|
||||
def test_easy_keyword(self):
|
||||
s = score_issue(issue(name="update docs and fix typo in readme"))
|
||||
assert s["friendliness"] > 50
|
||||
assert any("易上手" in sig for sig in s["signals"])
|
||||
|
||||
def test_hard_label(self):
|
||||
s = score_issue(issue(name="something", labels=["hard"]))
|
||||
assert "高难度标签" in str(s["signals"])
|
||||
|
||||
def test_difficulty_levels(self):
|
||||
easy = score_issue(issue(name="docs typo", labels=["good first issue"]))
|
||||
hard = score_issue(issue(name="refactor architecture performance concurrency"))
|
||||
assert easy["difficulty"] in ("入门", "较易")
|
||||
assert hard["difficulty"] in ("中等", "进阶")
|
||||
|
||||
def test_score_bounded(self):
|
||||
s = score_issue(issue(name="good first " * 10, labels=["good first issue", "beginner"]))
|
||||
assert 0 <= s["friendliness"] <= 100
|
||||
|
||||
def test_long_body_penalty(self):
|
||||
short = score_issue(issue(name="task", body="x" * 100))
|
||||
long = score_issue(issue(name="task", body="x" * 2000))
|
||||
assert long["friendliness"] <= short["friendliness"]
|
||||
|
||||
|
||||
class TestGuidance:
|
||||
def test_uses_number_when_available(self):
|
||||
g = build_guidance(score_issue(issue(name="fix", number=12)), "o", "r")
|
||||
assert "#12" in g
|
||||
|
||||
def test_uses_title_when_no_number(self):
|
||||
s = score_issue(issue(name="修复文档错别字", gid=999))
|
||||
g = build_guidance(s, "o", "r")
|
||||
# 无 web 序号时用标题引用,不出现 #999
|
||||
assert "#999" not in g
|
||||
assert "修复文档错别字" in g
|
||||
|
||||
def test_contains_fork_flow(self):
|
||||
g = build_guidance(score_issue(issue(name="task", number=1)), "Gitlink", "gitlink-cli")
|
||||
assert "repo +fork" in g
|
||||
assert "Gitlink" in g
|
||||
|
||||
|
||||
class TestBoard:
|
||||
def test_empty_candidates(self):
|
||||
result = {"owner": "o", "repo": "r", "total_issues": 3,
|
||||
"candidate_count": 0, "candidates": [], "all_scored": []}
|
||||
board = render_board(result, "o", "r")
|
||||
assert "暂未发现" in board
|
||||
|
||||
def test_board_with_candidates(self):
|
||||
cand = score_issue(issue(name="fix typo in docs", labels=["good first issue"], gid=100))
|
||||
result = {"owner": "o", "repo": "r", "total_issues": 5,
|
||||
"candidate_count": 1, "candidates": [cand], "all_scored": [cand]}
|
||||
board = render_board(result, "o", "r")
|
||||
assert "新手任务看板" in board
|
||||
assert "id:100" in board # 无 web 序号时标注全局 id
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
sys.exit(pytest.main([__file__, "-v"]))
|
||||
Loading…
Reference in New Issue