RAG/docs/code_retrieve_optimization.md

202 lines
4.6 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# 代码检索模块性能调优文档
## 1. 存储结构设计
### 1.1 核心设计原则
```
向量化文本: func_desc (函数描述,语义丰富)
元数据存储: func_body (完整函数代码,供问答使用)
```
### 1.2 存储格式
```python
Document(
text=func_desc, # 用于 embedding 的文本
metadata={
"func_body": "...", # 完整函数代码(核心)
"func_name": "...",
"file_path": "...",
"repo_id": "...",
# ... 其他元数据
}
)
```
## 2. 检索性能优化
### 2.1 Top-K 参数调优
| 场景 | 推荐 top_k | 说明 |
|------|-----------|------|
| 精确查找特定函数 | 3-5 | 减少噪声,提高精确度 |
| 探索性查询 | 10-15 | 获取更多上下文 |
| 复杂问题 | 20+ | 需要多个函数协作回答 |
### 2.2 元数据过滤
**优势**:在向量检索前过滤,减少计算量
```python
# 按仓库过滤
results = code_retrieve(query, repo_id="git_https___gitee_com_xxx")
# 按函数名精确查找
results = code_retrieve(query, func_name="push")
```
**性能提升**
- 无过滤:全库扫描 O(N)
- 有过滤:仅扫描子集 O(M)M << N
### 2.3 Embedding 模型选择
| 模型 | 维度 | 速度 | 适用场景 |
|------|------|------|----------|
| nomic-embed-text | 768 | | 通用代码检索 |
| qwen3-embedding | 768 | 中等 | 中文代码理解 |
| text-embedding-3 | 1536 | | 高精度需求 |
## 3. 检索质量优化
### 3.1 函数描述生成策略
当前实现
```python
def generate_func_desc(func_info):
parts = []
# 1. 函数类型和名称
if func_info.get("class_name"):
parts.append(f"{func_info['class_name']}类的{func_info['func_name']}方法")
else:
parts.append(f"{func_info['func_name']}函数")
# 2. 参数信息
if params:
parts.append(f"接收参数: {', '.join(param_str)}")
# 3. 返回值
if return_type:
parts.append(f"返回类型: {return_type}")
# 4. 文档字符串
if docstring:
parts.append(f"功能描述: {docstring}")
return ". ".join(parts)
```
**优化建议**
1. **添加函数调用关系**"调用 xx 函数"
2. **添加代码复杂度**"包含 xx 行代码"
3. **添加关键 API**"使用 requests "
### 3.2 相似度阈值
```python
# 过滤低质量结果
MIN_SIMILARITY_SCORE = 0.5
results = [r for r in results if r['score'] >= MIN_SIMILARITY_SCORE]
```
### 3.3 结果重排序
```python
def rerank_results(results, query):
"""
基于额外特征重排序:
1. 函数名匹配度
2. 文档字符串相关性
3. 代码长度适中度
"""
for r in results:
# 函数名完全匹配加分
if query.lower() in r['func_name'].lower():
r['score'] += 0.2
# 有文档字符串加分
if r['docstring']:
r['score'] += 0.1
return sorted(results, key=lambda x: x['score'], reverse=True)
```
## 4. 增量同步优化
### 4.1 变更检测
```python
def detect_changes(repo_path, last_sync_time):
"""
只检测变更的文件,避免全量扫描
"""
changed_files = []
for file in get_tracked_files(repo_path):
if get_file_mtime(file) > last_sync_time:
changed_files.append(file)
return changed_files
```
### 4.2 增量更新策略
1. **新增函数**直接添加
2. **修改函数**更新 embedding metadata
3. **删除函数**从向量库删除
## 5. 性能监控指标
### 5.1 关键指标
| 指标 | 目标值 | 监控方式 |
|------|--------|----------|
| 检索延迟 | < 500ms | 日志记录 |
| 准确率@5 | > 80% | 人工评估 |
| 覆盖率 | > 95% | 自动化测试 |
| 存储空间 | < 10GB | 系统监控 |
### 5.2 日志记录
```python
logger.info(f"[CodeRetrieve] 查询: '{query}', 找到 {len(results)} 个结果, 耗时 {elapsed_time:.2f}s")
```
## 6. 最佳实践
### 6.1 查询优化
**好的查询**
- "push 函数的功能是什么"
- "如何处理 GitHub issue"
- "webhook.py 中的 run 方法"
**差的查询**
- "代码"太宽泛
- "问题"无针对性
- "怎么写"不明确
### 6.2 索引维护
1. **定期重建索引**每月一次
2. **清理孤立文档**每周检查
3. **监控存储增长**设置告警
## 7. 故障排查
### 7.1 检索不到结果
1. 检查向量库是否为空
2. 检查 embedding 模型是否正常
3. 检查查询文本是否有效
### 7.2 结果质量差
1. 检查 func_desc 生成质量
2. 调整 top_k 参数
3. 考虑更换 embedding 模型
### 7.3 性能下降
1. 检查 ChromaDB 连接
2. 检查 embedding 服务负载
3. 考虑增加缓存层