Claude Skill

wechat-account-analyzer

公众号账号诊断工具是对任意公众号账号进行四维度量化评分(内容健康度、用户活跃度、内容核心数据、运营规范性),对标行业平均水平,输出可落地的运营优化建议。

LLM Mart · 0 points · 0 views 0 listing impressions 0 install-command copies
Virus-scanned Reviewed automatically before listing.

Full trust report

Download redfox-data-redfox-community-skills_wechat-account-analyzer-5e7b435.zip · 60 KB
Part of redfox-data/redfox-community — 66 skills

Install

skills CLI npx skills add https://github.com/redfox-data/redfox-community/tree/main/skills/wechat-account-analyzer
Claude Code claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install redfox-data-redfox-community@llmmart
Git git clone https://github.com/redfox-data/redfox-community.git

The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole redfox-data/redfox-community collection as a plugin from our marketplace. Git is the plain clone.

README

公众号账号诊断 / wechat-account-analyzer


简介

公众号账号诊断是一款智能分析工具,输入公众号名称即可获取四维度量化评分报告,对标行业平均水平,输出可落地的运营优化建议。

核心价值

  • 🎯 数据驱动决策:基于真实运营数据自动评分,告别主观判断,一次诊断即可看清账号全貌
  • 📊 四维全面诊断:内容健康度、用户活跃度、核心数据、运营规范性,覆盖运营关键维度
  • 💡 可执行建议:分紧急/重点/持续三级,每条建议含预期量化效果,直接指导优化动作
  • 🔍 对标行业水平:自动匹配同赛道账号,横向对比找差距,明确自身行业位置

适用对象

  • 👤 公众号号主 — 了解账号健康度,找到优化方向
  • 📱 新媒体运营 — 数据驱动运营策略,量化优化效果
  • 🏢 MCN 机构 — 评估账号价值,辅助签约决策
  • 🏷️ 品牌方与内容创作者 — 竞品分析,借鉴成功经验

功能特性

核心功能

  • 📊 四维度评分:内容健康度 / 用户活跃度 / 核心数据 / 运营规范性四维度诊断,v4.1 分类自适应(6 类账号自动调整权重/阈值),爆款率(10万+)显式计分
  • 🏆 智能评级:S/A/B/C/D/E 六级评级(S≥80 分)+ 行业对标分析,直观了解账号段位
  • 📈 作品数据展示:近期作品阅读量、点赞数、评论数、在看数一目了然
  • 💡 优化建议:按紧急程度分级,每条采用"问题 → 建议 → 预期"三段结构
  • 🔍 相似账号发现:自动匹配同赛道对标账号,横向对比找差距
  • 🔗 一键跳转:账号名称、作品标题均支持跳转微信查看详情

使用指南

直接用自然语言描述需求,无需记忆命令。

常用说法速查

意图 示例话术 效果
诊断公众号 "诊断十点读书" 输出完整诊断报告
多账号对比 "对比诊断 A 和 B" 横向对比 + 差异化建议
按 ID 查询 "诊断公众号ID xxx" 通过 ID 精确查询

输出示例

诊断报告以固定的 5 章节结构输出:

一、账号信息 → 名称、ID、认证主体、账号简介 二、综合评分 → 百分制评分 + 行业对标数据表 三、近期作品数据 → 最近作品详细数据表格 四、优化建议 → 紧急 / 重点 / 持续三级可执行建议 五、行业对标分析 → 对标结论 + 相似账号推荐


使用场景

场景 角色 示例问法 收益
号主自检优化 公众号号主 "诊断我的公众号" 明确账号健康度,量化优化效果
竞品对标分析 运营人员 "分析 XX 公众号" 掌握竞品动态,借鉴成功经验
MCN 批量评估 MCN 机构 "评估 XX 账号价值" 数据驱动签约决策
内容策略制定 内容创作者 "诊断并分析内容方向" 有据可依的内容规划

Skill manifest

公众号账号诊断

基于多维度数据分析,为公众号提供全面的账号健康诊断和优化建议


简介

公众号账号诊断 是一款智能分析工具,通过红狐 API 获取公众号真实运营数据,基于内容健康度、用户活跃度、核心数据表现、运营规范性四维度进行评分诊断,自动生成对标行业水平的优化建议。

  • 核心价值:告别主观判断,用数据驱动公众号运营决策。一次诊断即可看清账号全貌,明确优化方向
  • 适用对象:公众号号主、新媒体运营、MCN机构、品牌方、内容创作者
  • 技术基础:基于红狐API实时数据 + Python评分引擎 + Agent智能建议生成

功能特性

核心能力

功能 说明
📊 四维度评分 四维度诊断评分(内容健康度/用户活跃度/核心数据/运营规范性),v4.1分类自适应(6类账号自动调整权重/阈值),爆款率(10万+)显式计分
🏆 智能评级 S/A/B/C/D/E 六级评级(S≥80分)+ 行业对标分析
📈 数据可视化 近期作品数据表格,阅读/点赞/评论/在看一目了然
💡 优化建议 分紧急/重点/持续三级,每条建议含预期量化效果
🔍 相似账号 自动匹配同赛道对标账号,横向对比找差距
🔗 可跳转链接 账号名称、作品标题均支持一键跳转微信

特色亮点

  • ⚡ 一句话诊断:输入公众号名称即可获取完整报告,无需复杂配置
  • 🔄 多账号对比:支持同时诊断多个账号并生成对比报告
  • 📋 强制格式输出:5章节固定模板,章节/顺序/格式强制锁定,不得偏离

一键安装

前置条件

  • 已安装 Python 3.8+
  • 已注册 红狐Hub 账号并获取 API Key

安装步骤

  1. 将技能文件夹放入你的 Skills 目录
  2. 安装 Python 依赖:
    pip install requests
    
  3. 配置 API Key(见下方)

配置 API Key

获取 API Key

  1. 访问 红狐Hub 官网 了解服务详情
  2. 前往 注册页面 注册账号
  3. 新注册用户将获赠免费积分,可立即开始使用
  4. 注册登录后,在个人中心获取 API Key,格式为 ak_xxxxxxxx

设置环境变量

REDFOX_API_KEY 需从环境变量获取。若未设置,Agent 会主动帮你配置:

  • macOS/Linux:将 export REDFOX_API_KEY=<值> 追加到 ~/.zshrc 或 ~/.bashrc,然后 source 对应文件
  • Windows:使用 [Environment]::SetEnvironmentVariable("REDFOX_API_KEY", "<值>", "User") 设置用户级永久环境变量(需重启终端)

配置完成后验证:echo $REDFOX_API_KEY(macOS/Linux)或 echo %REDFOX_API_KEY%(Windows)


使用指南

触发方式

当用户提到以下任意关键词时自动激活:

  • "诊断公众号"、"账号分析"、"公众号体检"
  • "账号评估"、"查看XX公众号数据"、"分析XX账号"
  • 直接输入公众号名称 + "诊断/分析"(如"十点读书 诊断")

首次交互时 Agent 会主动打招呼并引导用户输入账号。

常用命令速查

命令 说明
诊断"公众号名称" 输入名称即可获取完整诊断报告
对比诊断"A"和"B" 多账号横向对比 + 差异化建议
诊断公众号ID"xxx" 通过公众号ID查询(非中文名)

使用示例

示例1:标准诊断

用户:诊断"十点读书"公众号 AI:输出完整诊断报告(账号信息→综合评分→作品数据→优化建议→行业对标)

示例2:多账号对比

用户:对比诊断"十点读书"和"洞见" AI:分别输出两个诊断报告 + 横向对比总结表格 + 差异化建议

示例3:未查到账号

用户:诊断“不存在的公众号123” AI:未查询到账号【不存在的公众号123】,该账号可能尚未收录或名称有误,请核实公众号名称后重试。

⚠️ 输出格式强制规范

诊断报告必须且仅能以5章节固定模板输出:一、账号信息 → 二、综合评分 → 三、近期作品数据 → 四、优化建议 → 五、行业对标分析。章节不可省略、顺序不可调换、格式不可偏离。所有字段名、表格结构、emoji 标识须与模板严格一致。违反模板的输出视为不合格。

详细技术规范、评分算法、输出格式模板请参见 核心工作流


使用场景

场景一:号主自检优化

需求:了解自己公众号在行业中的水平,找到优化方向

使用方式:

  1. 直接输入自己的公众号名称进行诊断
  2. 查看四维度评分找到短板
  3. 按优化建议的优先级逐步改进

预期收益:明确账号健康度,量化优化效果,避免盲目运营


场景二:竞品对标分析

需求:分析竞品账号的运营策略和数据表现

使用方式:

  1. 诊断目标竞品公众号
  2. 查看行业对标和相似账号
  3. 对比分析差异化优势与差距

预期收益:掌握竞品动态,借鉴成功经验,调整自身策略


场景三:MCN 批量评估

需求:评估旗下或潜在签约账号的价值

使用方式:

  1. 逐个诊断候选账号
  2. 横向对比综合评分和核心指标
  3. 结合优化建议评估成长潜力

预期收益:数据驱动签约决策,量化账号商业价值


场景四:内容策略制定

需求:根据账号诊断结果制定内容规划

使用方式:

  1. 诊断账号获取质量稳定性、内容垂直度等评分
  2. 根据"待优化模块"确定改进重点
  3. 参考行业基准调整内容方向和发布策略

预期收益:有据可依的内容规划,提升整体运营效率


项目架构

目录结构

公众号账号诊断/
├── SKILL.md                     # 技能说明文档(本文件)
├── scripts/
│   ├── wechat_analyzer.py       # 入口点:argparse 命令路由(query/sync/generate_html)
│   ├── api_client.py             # HTTP 通信层:红狐 API 凭证管理 + POST 请求封装
│   ├── scoring.py                # 评分引擎:作品辅助函数 + 四维度评分 + 等级判定
│   ├── analyzer.py               # 分析编排:单账号评分处理 + 查询/同步命令
│   ├── report.py                 # HTML 报告:模板替换 + 单账号/多账号报告生成
│   └── test_wechat_analyzer.py   # 单元测试:辅助函数 + 评分函数容错(42 个测试)
├── references/
│   ├── core_workflow.md         # 核心工作流(Agent执行参考)
│   ├── workflow_guide.md        # 工作流详细指南
│   └── api_guide.md             # API接口与评分逻辑说明
├── assets/
│   └── report_template.html     # 单账号HTML报告模板
└── output/                      # 输出目录(运行时自动生成)
    ├── raw_data.json            # API原始数据
    └── report_data.json         # 结构化诊断数据

技术栈

组件 技术
运行环境 Python 3.8+
HTTP 客户端 requests(原生)
数据源 红狐API (redfox.hk)
输出格式 Markdown + JSON + HTML
认证方式 Header X-API-KEY + 环境变量 REDFOX_API_KEY

核心模块

模块文件 函数 职责
wechat_analyzer.py main argparse 入口,路由到各子命令
api_client.py https_post 红狐 API HTTP POST 请求封装
api_client.py _get_credential 环境变量 → shell 配置文件读取 API Key
scoring.py _score_content_health 内容健康度(6子项加权,0-10分)
scoring.py _score_user_activity 用户活跃度(5子项加权,0-10分)
scoring.py _score_core_data 核心数据表现(6子项,0-31分)
scoring.py _score_operation_compliance 运营规范性(3子项,0-10分)
analyzer.py cmd_query API 查询 + 原始数据保存
analyzer.py _analyze_single_account 四维度评分计算 + 结构化输出
analyzer.py cmd_sync_notes 同步账号作品数据(订阅推送)
report.py cmd_generate_html 单账号 HTML 报告生成
report.py cmd_generate_multi_html 多账号对比 HTML 报告生成

常见问答

安装相关

Q1: 提示"未找到 REDFOX_API_KEY"怎么办?

A: 请按「一键安装 → 配置 API Key」中的步骤设置环境变量。如果不会配置,直接告诉 Agent,它会帮你操作。

Q2: 需要安装什么依赖?

A: 只需 requests 库:pip install requests。Python 3.8+ 即可运行。


使用相关

Q3: 支持按公众号ID查询吗?

A: 支持。如果知道公众号ID(如 gh_xxxxxxxx),Agent 会自动识别并使用 ID 查询。

Q4: 可以同时诊断多个账号吗?

A: 可以。输入"对比诊断 A 和 B"即可获得多账号对比报告。

Q5: 数据多久更新一次?

A: 数据来自红狐API,时效性取决于API缓存策略。报告末尾会标注数据置信度。


故障排除

Q6: 提示"未查询到该公众号信息"?

A: 可能原因:① 账号名称输入有误;② 该账号近期发文未达收录标准。请检查名称准确性后重试。

Q7: 数据看起来不对怎么办?

A: 所有数据来自红狐API接口,诊断结果仅反映接口返回的数据。如有疑问可联系红狐平台确认数据准确性。


获取帮助

Files (redfox-community)
  • assets
    • report_template.html 12 KB · in bundle
  • references
    • api_guide.md 17.5 KB
      # API接口与评分逻辑
      
      ## 接口请求
      
      ### 查询接口
      
      **请求方式**:POST
      **请求地址**:https://redfox.hk/story/api/gzh/data/searchUser
      
      | 参数名 | 类型 | 必填 | 说明 |
      |--------|------|------|------|
      | accountIds | Array<String> | 否 | 公众号ID列表(微信号/biz字段),与accountNames至少传一个 |
      | accountNames | Array<String> | 否 | 公众号名称列表,与accountIds至少传一个 |
      | source | String | 否 | 来源标识 |
      
      **请求体示例**:
      ```json
      {
        "accountNames": ["日食记"],
        "source": "公众号账号诊断-GitHub"
      }
      ```
      
      ### 同步接口(账号收录)
      
      **请求方式**:POST
      **请求地址**:https://redfox.hk/story/api/gzhUser/syncUserNotes
      
      | 参数名 | 类型 | 必填 | 说明 |
      |--------|------|------|------|
      | accountId | String | 是 | 公众号ID(微信号/biz字段) |
      | source | String | 否 | 来源标识 |
      
      ## 接口响应字段
      
      ### 账号数据(List格式)
      
      | 字段 | 说明 | 示例 |
      |------|------|------|
      | accountName | 公众号名称 | "十点读书" |
      | accountId | 公众号ID(微信号) | "gh_xxxxxxxx" |
      | avgReadCount | 平均阅读数 | 100000 |
      | avatar | 头像 | "https://..." |
      | description | 账号简介 | "聚焦天下大事..." |
      | verifyName | 账号认证信息 | "十点读书" |
      | accountType | 账号分类 | "新闻媒体" |
      | gmtCreate | 最近数据更新时间 | "2025-05-20 10:30:00" |
      
      ### 文章数据(works数组)
      
      | 字段 | 说明 | 示例 |
      |------|------|------|
      | id | 文章ID | 1616523288 |
      | title | 文章标题 | "中俄元首会晤成果文件清单" |
      | summary | 文章摘要 | "俄罗斯总统普京..." |
      | publishTime | 发布时间 | "2026-05-20 23:35:32" |
      | clicksCount | 阅读数 | 100001 |
      | likeCount | 点赞数 | 2126 |
      | commentCount | 评论数 | 9 |
      | watchCount | 在看数 | 622 |
      | shareCount | 分享数 | 3881 |
      | interactiveCount | 互动总数 | 102758 |
      | coverUrl | 封面图链接 | "https://mmbiz.qpic.cn/..." |
      | workUrl | 文章链接 | "https://mp.weixin.qq.com/s?..." |
      
      ## 🎯 四维度诊断体系概览
      
      > 详细的评分逻辑和阈值见下方「综合评分体系」章节。
      
      ### 综合评分 = 内容健康度×分类权重 + 用户活跃度×分类权重 + 内容核心数据×分类权重 + 运营规范性×分类权重
      
      | 维度 | 原始分 | general | news_politics | finance_opinion | emotion_lifestyle | 说明 |
      |------|--------|---------|---------------|-----------------|-------------------|------|
      | 内容健康度 | 0-10分 | 25% | 25% | 25% | 25% | 更新稳定性、内容垂直度、原创能力等 |
      | 用户活跃度 | 0-10分 | 20% | **8%** | 15% | **25%** | 互动率、留言质量等(含分类自适应阈值) |
      | 内容核心数据 | 0-31分 | 40% | **52%** | **45%** | 35% | 爆款产出力+阅读+点赞+评论+互动+时间 |
      | 运营规范性 | 0-10分 | 15% | 15% | 15% | 15% | 更新频率+发布时间+账号认证 |
      
      > **v4.1 分类自适应**:`_classify_account_category` 根据 description + 作品标题 + 互动结构自动识别账号分类(6类:时政新闻/财经理政/情感生活/知识教育/娱乐休闲/综合),调整权重、活跃度阈值、核心数据互动增益和阅读量保底值。
      
      ---
      
      ## 综合评分体系(100分制)
      
      **四维度加权均值(v4.1分类自适应):权重按账号分类动态调整,默认 25/20/40/15**
      
      ### 综合评分公式
      
      ```
      综合评分 = 内容健康度百分制×w1 + 用户活跃度百分制×w2 + 内容核心数据百分制×w3 + 运营规范性百分制×w4
      - v4.1: 权重按分类自适应(general=25/20/40/15, news_politics=25/8/52/15, finance_opinion=25/15/45/15, emotion_lifestyle=25/25/35/15, knowledge_edu=30/20/35/15, entertainment=20/15/45/20)
      ```
      
      ### 评分结构(v4.1 分类自适应,括号内为 general 默认值)
      
      | 评估模块 | 原始分 | 权重(general默认) | 说明 |
      |---------|--------|------------------|------|
      | **内容健康度** | 0-10分 | 25% | 6子项加权 |
      | **用户活跃度** | 0-10分 | 20% | 5子项加权(含分类自适应阈值) |
      | **内容核心数据** | 0-31分 | 40% | 爆款产出力(8)+阅读数(8)+点赞数(6)+评论数(4)+互动率(3)+发布时间(2) |
      | **运营规范性** | 0-10分 | 15% | 更新频率(5)+发布时间合理性(3)+账号认证(2) |
      
      ### 1. 内容健康度(0-10分原始分)
      
      **作为综合评分调整分输入之一**
      
      | 评估项 | 原始权重 | 评分标准(0-10分) |
      |--------|---------|-------------------|
      | 更新稳定性 | 15% | 10分=日更,8分=周更3-5次,5分=周更1-2次,<5=不规律 |
      | 内容垂直度 | 15% | 10分=极度垂直单一领域,7分=主领域+偶尔跨界,<5=内容杂乱 |
      | 原创能力 | 10% | 10分=90%+原创,8分=70-90%原创,5分=50-70%,<5=大量转载 |
      | 质量稳定性 | 10% | 阅读量波动系数(排除10万+截断值后计算),波动≤35%高分;全部10万+直接满分 |
      | 内容深度 | 5% | 长文比例、专业引用、独家观点 |
      | 形式创新 | 5% | 多媒体运用、互动形式、排版创新 |
      
      **计算公式:** 内容健康度 = 更新稳定性x0.15 + 内容垂直度x0.15 + 原创能力x0.10 + 质量稳定性x0.10 + 内容深度x0.05 + 形式创新x0.05
      
      ### 2. 用户活跃度(0-10分原始分)
      
      **作为综合评分调整分输入之一**
      
      | 评估项 | 原始权重 | 评分标准(0-10分) |
      |--------|---------|-------------------|
      | 互动率 | 20% | (点赞+在看+留言)/阅读量:>5%得高分,<1%低分 |
      | 留言质量 | 10% | 评论密度(公众号评论需审核,阈值远低于开放平台):≥0.05%=10分、≥0.02%=7、≥0.005%=4 |
      | 分享传播力 | 10% | 分享率:≥2%=10分、≥0.5%=7、≥0.2%=4(公众号2%+即为顶尖) |
      | 阅读完成率 | 5% | >60%高分,基于文章长度和行业基准 |
      | 活跃时段集中度 | 5% | 推送时间固定性,用户阅读时间规律 |
      
      **计算公式:** 用户活跃度 = 互动率x0.20 + 留言质量x0.10 + 分享传播力x0.10 + 阅读完成率x0.05 + 活跃时段x0.05
      
      > **v4.1 阅读量保底(分类感知)**:高均阅账号不应因低互动被过度惩罚。保底值按分类调整:时政号5万→≥70、3万→≥55;财经号5万→≥55、3万→≥45;通用5万→≥50、3万→≥40、1万→≥30。
      
      ### 3. 内容核心数据表现(0-31分原始分)
      
      | 评估项 | 分值 | 占比 | 说明 |
      |--------|------|------|------|
      | 爆款产出力 | 8分 | 25.8% | 10万+阅读文章占比(爆款率):≥40%=8分、≥30%=7、≥25%=6.5、≥20%=6、≥10%=4.5、≥5%=3 |
      | 阅读数表现 | 8分 | 25.8% | 内容传播广度:≥10万=8分、≥8万=7.5、≥5万=6.5、≥3万=5、≥2万=4、≥1万=3 |
      | 点赞数表现 | 6分 | 19.4% | 内容认可度:篇均≥3000或点赞率≥3%=6分、≥1000或≥1.5%=5、≥300或≥0.5%=3 |
      | 评论数表现 | 4分 | 12.9% | 用户参与度(微信评论需审核,阈值低于开放平台):篇均≥20或评论率≥0.05%=4分、≥10或≥0.03%=3、≥5或≥0.01%=2、≥2=1 |
      | 互动率表现 | 3分 | 9.7% | 综合互动质量:≥10%=3分、≥5%=2.5、≥3%=2、≥1.5%=1.5 |
      | 发布时间合理性 | 2分 | 6.5% | 发布时机优化:黄金时段≥60%=2分、≥30%=1.5、≥10%=1 |
      
      ### 4. 运营规范性(0-10分)
      
      **作为综合评分调整分输入之一**
      
      | 评估项 | 分值 | 说明 |
      |--------|------|------|
      | 更新频率 | 5分 | 7天发文>=5篇=5分,3-4篇=4分,2篇=3分,1篇=2分 |
      | 发布时间合理性 | 3分 | 固定时段发文>=60%=3分,>=40%=2分,>=20%=1分 |
      | 账号认证 | 2分 | 有认证=2分,无=0分 |
      
      ### 综合评级标准(100分制,基于公众号行业真实水平校准)
      
      | 得分区间 | 评级 | 等级说明 | 特征描述 |
      |---------|------|---------|---------|
      | **80-100分** | **标杆账号** | S级 | 均阅3万+、稳定日更、高互动的顶尖账号 |
      | **70-79分** | **优质账号** | A级 | 均阅1-3万、持续优质内容的优质账号 |
      | **60-69分** | **健康账号** | B级 | 正常运营、有稳定输出的健康账号 |
      | **50-59分** | **中等账号** | C级 | 基础运营待提升 |
      | **40-49分** | **亚健康账号** | D级 | 运营不稳定或互动偏低 |
      | **<40分** | **问题账号** | E级 | 多项指标不达标,急需全面优化 |
      
      ---
      
      ## 输出格式
      
      ### 诊断报告输出规范
      
      **严格按照以下5个章节顺序输出,不可省略任何章节。所有数值从脚本输出的JSON中对应字段取值,禁止自行编造。**
      
      #### JSON字段映射表
      
      脚本输出的JSON结构为 `{header, scores, content_health, user_activity, works, similar_accounts}`,报告各章节取值如下:
      
      | 报告位置 | 取值路径 | 说明 |
      |---------|---------|------|
      | 账号名称 | header.账号名 | |
      | 公众号ID | header.账号ID | |
      | 运营主体 | header.认证信息 | |
      | 账号简介 | header.账号简介 | |
      | 综合评分 | scores.综合评分 | 保留1位小数 |
      | 综合评级 | scores.综合评级 | 如"标杆账号" |
      | 综合等级 | scores.综合等级 | 如"S级" |
      | 近7天作品 | works | 数组,每项含标题(含超链接)/阅读数/点赞数/评论数/在看数/发布时间 |
      | 近7天作品提示 | works_hint | 无作品时为提示文字,否则为空
      | 行业对标 | scores.行业对标 | 含综合评分/平均阅读量/互动率/更新频率,每项有本账号/行业均值/头部账号 |
      | 相似账号 | similar_accounts | 数组,每项含账号名称/平均阅读数 |
      
      #### 一、账号信息(键值对,不使用表格)
      
      - **账号名称:** {header.账号名}
      - **公众号ID:** {header.账号ID}
      - **运营主体:** {header.认证信息}
      - **账号简介:** {header.账号简介}
      - **行业标签:** {header.账号标识}
      
      #### 二、综合评分
      
      **总分:{scores.综合评分} / 100分** — {scores.综合评级}({scores.综合等级})
      
      **行业对标表格:** 从scores.行业对标读取各指标的3个值(本账号/行业均值/头部账号),差距分析由智能体根据数据对比撰写。
      
      | 指标 | 本账号 | 行业均值 | 头部账号 | 差距分析 |
      |------|--------|---------|---------|---------|
      | **综合评分** | {scores.行业对标.综合评分.本账号} | {scores.行业对标.综合评分.行业均值} | {scores.行业对标.综合评分.头部账号} | [对比差距说明] |
      | **平均阅读量** | {scores.行业对标.平均阅读量.本账号} | {scores.行业对标.平均阅读量.行业均值} | {scores.行业对标.平均阅读量.头部账号} | [对比差距说明] |
      | **互动率** | {scores.行业对标.互动率.本账号} | {scores.行业对标.互动率.行业均值} | {scores.行业对标.互动率.头部账号} | [对比差距说明] |
      | **更新频率** | {scores.行业对标.更新频率.本账号} | {scores.行业对标.更新频率.行业均值} | {scores.行业对标.更新频率.头部账号} | [对比差距说明] |
      
      #### 三、近7天作品数据(最近5条)
      
      从works数组读取,每条作品输出6列。标题字段已包含Markdown超链接格式,直接原样输出。阅读数超过100000显示为"10w+",不足5条按实际条数输出,禁止增加其他列。
      
      > 若works_hint字段不为空,则不输出表格,仅输出works_hint中的提示文字。
      
      | 标题 | 阅读数 | 点赞数 | 评论数 | 在看数 | 发布时间 |
      |------|--------|--------|--------|--------|---------|
      | [示例文章标题](https://mp.weixin.qq.com/s/xxx) | 10w+ | 1280 | 356 | 89 | 05-22 08:30 |
      
      #### 四、优化建议
      
      分3个级别输出,基于各维度评分数据生成具体可执行建议:
      
      **紧急优化(1-2项)**
      1. **[具体问题]** → [可执行建议](预期:[量化效果])
      
      **重点优化(2-3项)**
      1. **[具体问题]** → [可执行建议](预期:[量化效果])
      
      **持续优化(1-2项)**
      1. **[优化方向]** → [策略建议](预期:[长期收益])
      
      #### 五、行业对标分析
      
      **对标结论:**
      - **综合评级:{scores.综合评级图标} {scores.综合等级}({scores.综合评分}分)**,[行业地位描述]
      - **核心优势:** [基于评分数据分析列出2-3个核心优势]
      - **与头部差距:** [说明与头部账号的差距]
      - **优化方向:** [基于评分数据分析列出2-3个优化方向]
      
      **相似账号:** 从similar_accounts数组读取,最多5条。
      
      | 序号 | 账号名称 | 平均阅读 |
      |------|---------|---------|---------|
      | 1 | [similar_accounts[0].账号名称] | [similar_accounts[0].平均阅读数] |
      | ... | ... | ... | ... |
      
      ---
      
      *诊断时间:[YYYY-MM-DD] | 数据来源:红狐API接口 | 置信度:[高/中] | 阅读数>10万显示为10w+ | 禁止输出西瓜指数*
      
      ---
      
      ## ⚠️ 注意事项
      
      ### 数据来源限制(重要)
      - 🔒 **接口数据唯一性:** 所有账号信息(基础信息、运营数据等)和内容文章数据(阅读数、点赞数、评论数等)**必须通过API接口获取**
      - ❌ **禁止自主采集:** 严禁从其他渠道(如新榜、清博、微信后台、第三方平台等)自主获取或估算数据
      - ✅ **接口优先原则:** 所有诊断分析必须基于接口返回的真实数据,不得使用估算值或推测值
      - 📊 **数据准确性:** 接口返回的数据为唯一可信数据源,所有评分和诊断结论必须基于此
      - 📖 **阅读数显示规则:** 阅读数超过100000时显示为"10w+"而非具体数字
      
      ### 数据准确性
      - 📊 **接口数据:** 所有数据来自红狐API接口,确保权威性和准确性
      - ⏰ **时效性:** 标注数据采集时间,提醒用户数据可能随时间变化
      - 🔍 **置信度标注:** 在报告末尾标注数据置信度(高=接口实时数据,中=接口缓存数据)
      - ⚠️ **未查询到账号:** 当接口返回空结果时,明确告知用户"未查询到该公众号信息",并提示"请输入准确的公众号名称或公众号ID重新查询",**严禁生成估算报告**
      
      ### 诊断客观性
      - ✅ **优势+问题并重:** 既要指出问题,也要肯定优势,避免片面评价
      - 📝 **数据支撑:** 所有诊断结论需有数据或现象支撑,不可主观臆断
      - 🎯 **可操作建议:** 建议必须具体可执行,包含预期效果,避免空泛表述
      - ⚖️ **评分一致性:** 同类型账号评分标准保持一致,确保横向可比
      
      ### 特殊场景处理
      - ❌ **未查询到账号:** 当无法找到用户输入的公众号信息,或**没有查询到公众号ID**时,明确告知用户"未查询到该公众号信息",并提示"请输入准确的公众号名称或公众号ID重新查询",**严禁生成估算报告**
      - 🔍 **新账号(<6个月):** 数据样本不足,降低诊断置信度至"中"或"低",侧重内容方向评估
      - 📉 **停更账号:** 标注"账号已停更",基于历史数据分析,提示数据时效性
      - 🎭 **矩阵账号:** 如属MCN矩阵,说明团队协作特点,评估需考虑资源协同效应
      - 💼 **企业号vs个人号:** 企业号侧重品牌传播和客户服务,个人号侧重IP打造和内容质量
      - 📰 **媒体号vs自媒体:** 媒体号有采编团队,标准更高;自媒体侧重个人特色和用户粘性
      
      ### 输出规范
      - 输出综合评分总分(100分制),不输出评分维度表格、优势模块和待优化模块
      - 🚫 **禁止输出内容健康度明细分析**,不得出现"内容健康度分析"独立章节或子项评分表格
      - 🚫 **禁止输出用户活跃度明细分析**,不得出现"用户活跃度分析"独立章节或子项评分表格
      - 📊 评分需精确到小数点后1位(如7.3/10)
      - 💡 优化建议按优先级分3级(紧急/重点/持续),数量分别为1-2项、2-3项、1-2项
      - 📈 行业对标需包含4个核心指标:综合评分、平均阅读量、互动率、更新频率
      - 📝 对标结论需包含:综合评级、核心优势、与头部差距、优化方向
      - 🎨 严格遵循Markdown格式,确保表格对齐、标题层级正确
      - 🔒 **所有数据必须来自接口**,不得使用估算值或外部数据源
      
      ### 常见行业基准参考(2024-2026)
      | 账号类型 | 平均阅读量 | 打开率 | 互动率 | 更新频率 |
      |---------|-----------|--------|--------|---------|
      | 头部大号 | 5万+ | 3-5% | 5-8% | 日更 |
      | 中腰部账号 | 5000-3万 | 2-4% | 3-5% | 周更3-5次 |
      | 小账号 | 500-3000 | 1-3% | 2-4% | 周更1-3次 |
      | B端专业号 | 2000-1万 | 5-10% | 2-3% | 周更2-4次 |
      | C端娱乐号 | 3000-5万 | 2-5% | 5-10% | 日更或隔日更 |
      
      ---
      
      ## 📚 使用示例
      
      ### 示例1:完整诊断(标准场景)
      ```
      用户:诊断"十点读书"公众号
      AI:[输出完整诊断报告,包含基本信息、核心数据、四维度诊断、优化建议、行业对标]
      ```
      
      ### 示例2:快速评估(简化场景)
      ```
      用户:分析一下"咪蒙"这个号
      AI:[输出完整诊断报告,可适当简化数据表格,但必须包含四维度诊断和优化建议]
      ```
      
      ### 示例3:对比诊断(多账号场景)
      ```
      用户:对比诊断"十点读书"和"洞见"
      AI:[分别输出两个账号的完整诊断报告] + [横向对比总结表格] + [差异化建议]
      ```
      
      ### 示例4:使用公众号ID查询
      ```
      用户:诊断公众号ID为"isdushu"的账号
      AI:[先确认账号名称,再输出完整诊断报告]
      ```
      
      ### 示例5:特殊场景处理
      ```
      用户:帮我看看"XX公众号"怎么样
      AI:[检测到新账号/停更账号等特殊情况,在报告中标注置信度和特殊说明]
      ```
      
      ### 示例6:未查询到账号
      ```
      用户:诊断"不存在的公众号123"
      AI:抱歉,未查询到"不存在的公众号123"的公众号信息。请输入准确的公众号名称或公众号ID重新查询。
      ```
    • core_workflow.md 18 KB
      # 公众号诊断核心工作流
      
      > Agent 执行诊断任务时的完整技术参考,包含 API、评分算法、输出规范和行为约束。
      
      ---
      
      ## 1. 工作流步骤
      
      ### 步骤1:识别输入
      - 用户提供公众号名称(中文)或公众号ID(纯数字/字母数字组合)
      - 中文名称 → 使用 `--account_names` 参数查询
      - 纯数字/字母数字组合 → 使用 `--account_ids` 参数查询
      
      **首次交互开场白**:
      > 你好!我是公众号账号诊断宗师,可以帮你深度拆解公众号账号的运营数据和商业价值。
      > 请提供你想分析的公众号ID或公众号昵称均可。
      
      ### 步骤2:查询数据
      
      按名称查询:
      ```bash
      python scripts/wechat_analyzer.py query --account_names "公众号昵称"
      ```
      
      按ID查询:
      ```bash
      python scripts/wechat_analyzer.py query --account_ids "账号ID"
      ```
      
      - 脚本自动保存 `output/raw_data.json` 和 `output/report_data.json`
      - 脚本输出结构化 JSON,包含 `status`、`query_type`、`message`
      
      ### 步骤3:处理查询结果
      
      **成功时**(`status: "success"`, `query_type: "single"`):
      - 读取 `output/report_data.json`,按第4节输出格式生成诊断报告
      
      **未找到时**(`status: "success"`, `query_type: "not_found"`):
      - 若返回 `candidates` 数组(有相近账号),将候选列表展示给用户(含账号名称、公众号ID `wxId`、微信号 `account`),引导用户确认正确名称或用微信号/ID重试
      - 若无候选账号,告知用户该账号尚未收录或名称有误
      - **严禁生成估算报告**
      
      **错误时**(`status: "error"`):
      - 检查错误信息,如为凭证问题则引导配置 `REDFOX_API_KEY`
      
      ### 步骤4:生成报告
      - **不生成 HTML 文件**,直接读取 `output/report_data.json`,按第4节格式在对话中输出文字诊断报告
      - 严格按照第4节输出格式的5个章节顺序输出
      - 所有数值从 `report_data.json` 对应字段取值,禁止自行编造
      
      ---
      
      ## 2. API 调用配置
      
      ### 接口地址(双接口查询链路)
      
      | 步骤 | 用途 | 地址 |
      |------|------|------|
      | 第一步 | 按名称搜索账号,获取微信号 | `https://redfox.hk/story/api/gzh/data/searchUser` |
      | 第二步 | 按微信号精确查询完整数据 | `https://redfox.hk/story/api/gzhUser/queryData` |
      
      ### 请求配置
      
      - **请求方式**:POST
      - **请求头**:`Content-Type: application/json` + `X-API-KEY: {REDFOX_API_KEY}`
      - **认证方式**:Header `X-API-KEY`,值从环境变量 `REDFOX_API_KEY` 获取
      - **成功响应码**:统一为 `2000`(非 HTTP 标准 200)
      
      ### 接口一:searchUser(搜索账号)
      
      **请求体**:
      ```json
      { "keyword": "公众号名称", "offset": 0 }
      ```
      
      **关键返回字段**:
      
      | 字段 | 类型 | 说明 |
      |------|------|------|
      | accountName | String | 公众号名称 |
      | account | String | **微信号**(如 `duhaoshu`),作为 queryData.accountIds 参数 |
      | wxId | String | 公众号 AppID(`gh_xxx` 格式) |
      | verifyName | String | 认证信息 |
      | description | String | 账号简介 |
      | accountType | String | 账号分类 |
      
      > ⚠️ 名称匹配必须严格相等(`accountName == 输入名称`),匹配失败时返回候选列表(含 account + wxId)供用户选择。
      
      ### 接口二:queryData(精确查询完整数据)
      
      **请求体**:
      ```json
      {
        "accountIds": ["duhaoshu"],   // 微信号(account 字段),非 gh_xxx
        "accountNames": ["十点读书"]
      }
      ```
      
      **关键返回字段**:
      
      | 字段 | 类型 | 说明 |
      |------|------|------|
      | accountId | String | 公众号 AppID(`gh_xxx`) |
      | accountName | String | 公众号名称 |
      | verifyName | String | 认证信息 |
      | description | String | 账号简介 |
      | works | Array | 近期作品列表(最多20条) |
      | similarAccounts | Array | 相似账号列表 |
      
      **works 子字段**:
      
      | 字段 | 类型 | 说明 |
      |------|------|------|
      | title | String | 文章标题 |
      | workUrl | String | 文章链接(完整 mp.weixin.qq.com URL) |
      | clicksCount | Integer | 阅读数(主字段) |
      | likeCount | Integer | 点赞数 |
      | commentCount | Integer | 评论数 |
      | shareCount | Integer | 分享数 |
      | watchCount | Integer | 在看数 |
      | isOriginal | Boolean | 是否原创 |
      | publishTime | Long | 发布时间(Unix 毫秒时间戳) |
      
      ---
      
      ## 3. 评分体系
      
      ### 3.1 四维度诊断体系(v4.1 分类自适应:综合评分 = 内容健康度×分类权重 + 用户活跃度×分类权重 + 内容核心数据×分类权重 + 运营规范性×分类权重)
      
      > **v4.1 分类自适应**:根据账号分类(时政新闻/财经理政/情感生活/知识教育/娱乐休闲/综合)自动调整四维度权重、用户活跃度阈值、核心数据互动增益和阅读量保底值。分类由 `_classify_account_category` 根据 description + 作品标题 + 互动结构三级信号融合自动识别。
      
      #### 维度一:内容健康度(综合评分权重 25%)
      
      | 评估项 | 权重 | 评分标准(0-10分) |
      |--------|------|-------------------|
      | 更新稳定性 | 15% | 10分=日更,8分=周更3-5次,5分=周更1-2次,<5=不规律 |
      | 内容垂直度 | 15% | 10分=极度垂直单一领域,7分=主领域+偶尔跨界,<5=内容杂乱 |
      | 原创能力 | 10% | 10分=90%+原创,8分=70-90%原创,5分=50-70%,<5=大量转载 |
      | 质量稳定性 | 10% | 阅读量波动系数(排除10万+截断值后计算),波动≤35%得高分;全部10万+直接满分 |
      | 内容深度 | 5% | 长文比例、专业引用、独家观点 |
      | 形式创新 | 5% | 多媒体运用、互动形式、排版创新 |
      
      ```
      内容健康度 = 更新稳定性×0.15 + 内容垂直度×0.15 + 原创能力×0.10 + 质量稳定性×0.10 + 内容深度×0.05 + 形式创新×0.05
      ```
      
      #### 维度二:用户活跃度(综合评分权重 20%)
      
      | 评估项 | 权重 | 评分标准(0-10分) |
      |--------|------|-------------------|
      | 互动率 | 20% | (点赞+在看+留言)/阅读量:>5%得高分,<1%低分 |
      | 留言质量 | 10% | 评论密度(公众号评论需审核):≥0.05%=10分、≥0.02%=7、≥0.005%=4 |
      | 分享传播力 | 10% | 分享率:≥2%=10分、≥0.5%=7、≥0.2%=4 |
      | 阅读完成率 | 5% | 基于文章长度和行业基准,>60%高分 |
      | 活跃时段集中度 | 5% | 推送时间固定性,用户阅读时间规律 |
      
      > **v4.1 阅读量保底**:高均阅账号不应因低互动被过度惩罚。保底值按分类调整:时政号均阅≥5万→≥70、≥3万→≥55;财经号≥5万→≥55、≥3万→≥45;通用≥5万→≥50、≥3万→≥40、≥1万→≥30。
      >
      > **v4.1 核心数据互动增益**:时政号互动类子项(点赞/评论/互动率)得分×2.0(封顶),财经号×1.5(封顶)。因时政/新闻号读者被动消费内容,互动天然低但传播价值极高。
      
      ```
      用户活跃度 = 互动率×0.20 + 留言质量×0.10 + 分享传播力×0.10 + 阅读完成率×0.05 + 活跃时段×0.05
      ```
      
      ### 3.2 综合评分体系(100分制,四维度加权)
      
      **综合评分 = 四维度加权均值**
      
      ```
      综合评分 = 内容健康度百分制×w1 + 用户活跃度百分制×w2 + 内容核心数据百分制×w3 + 运营规范性百分制×w4
      - v4.1 权重(w1/w2/w3/w4)按分类自适应:general=25/20/40/15, news_politics=25/8/52/15, finance_opinion=25/15/45/15, emotion_lifestyle=25/25/35/15, knowledge_edu=30/20/35/15, entertainment=20/15/45/20
      ```
      
      四维度各自独立评分后按权重加权汇总为100分制综合评分,其中内容核心数据维度包含爆款产出力(10万+率)等6项子指标,确保高影响力账号获得合理高分。
      
      **各维度原始评分**:
      
      | 评估模块 | 原始分 | 说明 |
      |---------|--------|------|
      | 内容健康度 | 0-10分 | 6子项加权 |
      | 用户活跃度 | 0-10分 | 5子项加权 |
      | 内容核心数据 | 0-31分 | 爆款产出力(8)+阅读数(8)+点赞数(6)+评论数(4)+互动率(3)+发布时间(2) |
      | 运营规范性 | 0-10分 | 更新频率(5)+发布时间合理性(3)+账号认证(2) |
      
      **核心数据子项分档**:
      - 爆款产出力(10万+率):≥40%=8分、≥30%=7、≥25%=6.5、≥20%=6、≥10%=4.5、≥5%=3、≥2%=2、>0=1、0=0
      - 质量稳定性计算时排除10万+截断值(100001),避免爆款账号被波动惩罚;全部10万+直接满分
      
      ### 3.3 综合评级标准(基于公众号行业真实水平校准)
      
      | 得分区间 | 评级 | 等级说明 | 特征描述 |
      |---------|------|---------|---------|
      | **80-100分** | 🏆 标杆账号 | S级 | 均阅3万+、稳定日更、高互动的顶尖账号 |
      | **70-79分** | ⭐ 优质账号 | A级 | 均阅1-3万、持续优质内容的优质账号 |
      | **60-69分** | ✅ 健康账号 | B级 | 正常运营、有稳定输出的健康账号 |
      | **50-59分** | 📊 中等账号 | C级 | 基础运营待提升 |
      | **40-49分** | ⚠️ 亚健康账号 | D级 | 运营不稳定或互动偏低,需调整改进 |
      | **<40分** | ❌ 问题账号 | E级 | 多项指标不达标,急需全面诊断 |
      
      ---
      
      ## 4. 诊断报告输出格式(强制规范)
      
      > 🔴 **铁律:必须逐章逐节按以下模板输出,章节不可省略、顺序不可调换、格式不可偏离、字段名不可改写。所有数值从 `output/report_data.json` 中对应字段取值,禁止自行编造。凡不按此格式输出的诊断报告均视为不合格。**
      
      ### JSON 数据结构
      
      脚本输出的 JSON 结构为 `{header, scores, content_health, user_activity, works, similar_accounts}`。
      
      ### JSON 字段映射表
      
      | 报告位置 | 取值路径 | 说明 |
      |---------|---------|------|
      | 账号名称 | header.账号名 | |
      | 公众号ID | header.账号ID | |
      | 账号链接 | header.账号链接 | `https://open.weixin.qq.com/qr/code?username={账号ID}` |
      | 运营主体 | header.认证信息 | |
      | 账号简介 | header.账号简介 | |
      | 综合评分 | scores.综合评分 | 保留1位小数 |
      | 综合评级 | scores.综合评级 | 如"标杆账号" |
      | 综合等级 | scores.综合等级 | 如"S级" |
      | 近7天作品 | works | 数组,每项含标题(含超链接)/阅读数/点赞数/评论数/在看数/发布时间 |
      | 近7天作品提示 | works_hint | 无作品时为提示文字,否则为空
      | 行业对标 | scores.行业对标 | 含5项指标每项3个值(本账号/行业均值/头部账号) |
      | 相似账号 | similar_accounts | 数组,每项含账号名称(可点击跳转),链接格式 `https://open.weixin.qq.com/qr/code?username={账号ID}` |
      
      ---
      
      ### 一、账号信息
      
      **使用表格格式输出:**
      
      | 项目 | 内容 |
      |------|------|
      | 账号名称 | [{header.账号名}]({header.账号链接}) |
      | 公众号ID | {header.账号ID} |
      | 运营主体 | {header.认证信息},为空时填"未认证" |
      | 账号简介 | {header.账号简介},为空时填"无" |
      | 账号类型 | {header.账号类型},为空时填"未知" |
      
      ### 二、综合评分
      
      **整体评级:「{scores.综合评分}」 / 100分 — {scores.综合评级}({scores.综合等级})**
      
      > 综合评分保留1位小数,使用中文书名号包裹评分值。
      
      **核心指标对标表**(差距分析由 Agent 根据数据对比撰写,优于行业用 🔺 前缀,低于行业用 🔻 前缀,持平用 ➖ 前缀):
      
      | 指标 | 本账号 | 行业均值 | 头部账号 | 差距分析 |
      |------|--------|---------|---------|---------|
      | **综合评分** | {scores.行业对标.综合评分.本账号} | {scores.行业对标.综合评分.行业均值} | {scores.行业对标.综合评分.头部账号} | [差距分析] |
      | **平均阅读量** | {scores.行业对标.平均阅读量.本账号} | {scores.行业对标.平均阅读量.行业均值} | {scores.行业对标.平均阅读量.头部账号} | [差距分析] |
      | **互动率** | {scores.行业对标.互动率.本账号} | {scores.行业对标.互动率.行业均值} | {scores.行业对标.互动率.头部账号} | [差距分析] |
      | **更新频率** | {scores.行业对标.更新频率.本账号} | {scores.行业对标.更新频率.行业均值} | {scores.行业对标.更新频率.头部账号} | [差距分析] |
      
      ### 三、近7天作品数据(最近5条)
      
      从 works 数组读取,标题含 Markdown 超链接直接输出。
      
      > 🔴 **链接禁止截断除内容:`标题`字段已包含完整的 Markdown 链接格式 `[标题文字](URL)`,输出时必须直接使用该字段完整内容,禁止用 `(…)` 或任意内容替换 URL 部分;可适度截短显示的标题文字,但完整链接地址必须保留。**
      
      若 `works_hint` 不为空,不输出表格,以引用块输出提示文字。
      
      若作品数据中存在明显异常(如大量0阅读但发布时间为近期),在表格后添加 ⚠️ 注意说明。
      
      | 标题 | 阅读数 | 点赞数 | 评论数 | 在看数 | 发布时间 |
      |------|--------|--------|--------|--------|---------|
      | ... | ... | ... | ... | ... | ... |
      
      ### 四、优化建议
      
      分3级,使用 emoji 标识优先级。每项包含:**问题** → **建议** → **预期** 三段结构:
      
      **🔺 紧急优化(1-2项)**
      1. **{具体问题}**
         - 问题:[问题描述]
         - 建议:[可执行建议]
         - 预期:[量化效果]
      
      **📌 重点优化(2-3项)**
      1. **{具体问题}**
         - 问题:[问题描述]
         - 建议:[可执行建议]
         - 预期:[量化效果]
      
      **✅ 持续优化(1-2项)**
      1. **{优化方向}**
         - 建议:[策略建议]
         - 预期:[长期收益]
      
      ### 五、行业对标分析
      
      **对标结论**采用分段结构,包含:综合评级、核心优势(2-3个)、与头部差距、优化方向(2-3个):
      
      > **综合评级**:[具体描述]
      > **核心优势**:
      > - [优势1]
      > - [优势2]
      > **与头部差距**:[差距描述]
      > **优化方向**:
      > - [方向1]
      > - [方向2]
      
      **相似账号**:从 similar_accounts 读取,最多5条。**账号名称必须添加可点击跳转链接**,格式为 `[{similar_accounts[i].账号名称}](https://open.weixin.qq.com/qr/code?username={similar_accounts[i].账号ID})`。如发现账号矩阵等有趣现象,在表格后追加“💨 有意思的发现”段落。
      
      | 序号 | 账号名称 |
      |------|---------|
      | 1 | [{similar_accounts[0].账号名称}](https://open.weixin.qq.com/qr/code?username={similar_accounts[0].账号ID}) |
      | ... | ... |
      
      ---
      
      **诊断报告末尾必须依次添加以下两段提醒**:
      
      > *诊断时间:[YYYY-MM-DD] | 数据来源:[红狐数据(redfox.hk)](https://redfox.hk/) 更多平台数据、优质数据可前往查看API接口 | 置信度:[高/中/低] | 阅读数>10万显示为10w+*
      
      > 💡 **订阅提醒**:如需持续追踪该账号数据变化,可回复"订阅 {账号名称}",系统将自动同步最新作品数据(约需30分钟)。首次订阅后可解锁完整7天作品数据。
      
      ---
      
      ## 5. Agent 行为规范(强制)
      
      ### 🔴 最高优先级:输出格式铁律
      - **诊断报告必须严格按照第4节定义的5章节模板输出,一字不漏**
      - 章节顺序:一→二→三→四→五,不可调换、不可省略
      - 每个章节内部的格式(表格/列表/引用块/emoji)**必须**与模板一致
      - 所有字段名从 `report_data.json` 取值,**禁止改写字段名**(如不可将"整体评级"写成"综合得分")
      - 凡不遵守以上规则的输出,视为不合格,必须重新生成
      
      ### 数据来源限制
      - 🔒 所有账号信息和文章数据**必须通过API接口获取**
      - ❌ 严禁从其他渠道自主获取或估算数据
      - ✅ 所有诊断分析必须基于接口返回的真实数据
      
      ### 未查询到账号
      - 告知用户未查询到,并输出以下三条可能原因:
        1. **ID 格式问题**:`gh_xxx` 是公众号 AppID 格式,如果知道微信号(如 `duhaoshu` 这类形式),可尝试用微信号查询
        2. **尚未收录**:该账号发文量较少,未被红狐数据库收录
        3. **申请收录**:收录到红狐后可进行每日定时数据追踪与数据分析,如需收录请发送邮件至 redfoxdata@proton.me
      - **严禁生成估算报告**
      
      ### 特殊场景处理
      - **新账号(<6个月)**:数据样本不足,置信度降至"中"或"低",侧重内容方向评估
      - **停更账号**:标注"账号已停更",提示数据时效性
      - **矩阵账号**:说明团队协作特点,评估考虑资源协同效应
      - **企业号 vs 个人号**:企业号侧重品牌传播,个人号侧重IP打造
      - **媒体号 vs 自媒体**:媒体号标准更高,自媒体侧重个人特色
      
      ### 输出禁止项
      - 🚫 禁止输出内容健康度明细分析独立章节
      - 🚫 禁止输出用户活跃度明细分析独立章节
      - 🚫 禁止输出西瓜指数相关字段或数据
      - 🚫 禁止输出评分维度表格、优势模块和待优化模块
      - 🚫 禁止在综合评分行使用"总分"字样,统一使用"整体评级"
      
      ### 输出规范
      - 综合评分总分100分制,精确到小数点后1位,使用中文书名号包裹如「92.6」
      - 优化建议按优先级分3级,使用 emoji 标识:🔺 紧急优化、📌 重点优化、✅ 持续优化
      - 优化建议每项采用"问题→建议→预期"三段结构
      - 行业对标含4个核心指标:综合评分、平均阅读量、互动率、更新频率
      - 行业对标差距分析使用 emoji 前缀:🔺 优于行业、🔻 低于行业、➖ 持平
      - 阅读数>10万显示为"10w+"
      - 账号信息使用表格格式(项目 | 内容),不用键值对列表
      - 作品数据存在异常时(如刚发布阅读为0),在表格后加 ⚠️ 注意说明
      - 相似账号分析中发现矩阵/品牌关联等有趣现象,追加"💨 有意思的发现"段落
      - 相似账号名称必须添加可点击跳转链接
      - 报告末尾必须追加 💡 订阅提醒(引导用户回复"订阅 {账号名称}")
      - 空值字段处理:认证信息为空填"未认证",简介为空填"无"
      
      ### 行业基准参考(基于公众号真实行业水平)
      
      | 账号类型 | 平均阅读量 | 互动率 | 更新频率 | 备注 |
      |---------|-----------|--------|---------|------|
      | 行业标杆(S级) | 3-10万 | 10-25% | 日更/隔日更 | 极少账号能达到 |
      | 头部优质(A级) | 1-3万 | 5-10% | 周更3-5次 | 行业前5% |
      | 中腰部(B级) | 5000-1万 | 3-5% | 周更2-3次 | 行业前20% |
      | 普通账号(C级) | 2000-5000 | 1-3% | 周更1-2次 | 行业平均水平 |
      | 小账号(D级) | 500-2000 | <1% | 不定期 | 待运营提升 |
      
    • workflow_guide.md 22.3 KB
      # 公众号账号分析工作流指南
      
      ## 目录
      1. [整体流程](#整体流程)
      2. [评分体系](#评分体系)
      3. [数据获取与脚本调用](#数据获取与脚本调用)
      4. [各模块分析指南](#各模块分析指南)
      5. [水平衡量基准数据](#水平衡量基准数据)
      6. [多账号对比](#多账号对比)
      
      ## 整体流程
      
      ### 步骤1:获取公众号信息
      - 用户可以提供公众号ID或公众号昵称,两种方式均可查询
      - 公众号ID:如 `gh_xxxxxxxx`、纯数字或字母数字组合
      - 公众号昵称:如 `十点读书`、`视觉志` 等中文名称
      
      **⚠️ 输入识别规则**:
      - 用户输入为纯数字或字母数字组合时,识别为公众号ID,使用 `--account_ids` 参数查询
      - 用户输入为中文名称时,识别为公众号昵称,使用 `--account_names` 参数查询
      
      **标准开场白**(首次交互时输出):
      "你好!我是公众号账号诊断宗师,可以帮你深度拆解公众号账号的运营数据和商业价值。
      请提供你想分析的公众号ID或公众号昵称均可。"
      
      ### 步骤2:查询数据
      按ID查询:
      ```bash
      python scripts/wechat_analyzer.py query --account_ids "账号ID"
      ```
      按昵称查询:
      ```bash
      python scripts/wechat_analyzer.py query --account_names "公众号昵称"
      ```
      - 脚本自动保存 `output/raw_data.json`
      - 输出结构化JSON结果,包含评分和指标
      
      **⚠️ 账号数据为空处理(未查询到账号)**:
      - 如果脚本返回 `status: "error"` 且提示未查询到账号,说明该账号不在数据库中
      - **立即停止后续步骤**,执行以下处理:
        1. 调用数据同步接口:
           ```bash
           python scripts/wechat_analyzer.py sync_notes --red_ids "用户输入的账号ID"
           ```
           - 同步接口:`https://redfox.hk/story/api/gzh/data/queryWorkList`,参数 `redId` 从用户输入中获取
        2. 输出提示:
           ```
           非常抱歉,未查询到该账号。请检查账号id是否有误,若输入准确,可能当前账号近期发文数据未达到收录标准。
           我们将为你实时去采集数据,30min后向您推送新的诊断报告,请问是否接收?
           1. 接收推送
           2. 暂不接收
           ```
        3. 用户输入编号"1"或回复"接收"后,执行订阅流程:
           a. **【必须】创建日程任务**(使用calendar_create工具):
              - summary: 推送xxx账号诊断报告
              - dtstart: 当前时间+30分钟(格式YYYYMMDDHHMM)
              - description: 重新查询账号数据并推送诊断报告(python scripts/wechat_analyzer.py query --account_ids <用户输入的账号ID>)
           b. 输出提示:「订阅成功!系统将在30分钟后自动向您推送该账号的诊断报告,请耐心等待」
        4. 用户输入编号"2"或回复"暂不接收",结束本次分析
      
      **⚠️ 作品数据为空处理(直接进入订阅流程)**:
      - 如果查询到账号但 `works` 为空数组,脚本返回 `query_type: "need_sync"` 和 `need_sync` 字段
      - **立即停止后续步骤**,不输出分析报告,直接进入订阅流程:
        1. 输出提示(xxx为账号昵称):
           ```
           为您获取作品中,30分钟后会更新数据后可向您推荐,请问是否订阅"xxx"账号?
           1. 订阅账号xxx
           2. 暂不订阅
           ```
        2. 用户输入编号"1"或回复"订阅"后,执行订阅流程:
           a. **【必须】创建日程任务**(使用calendar_create工具):
              - summary: 推送xxx账号诊断报告
              - dtstart: 当前时间+30分钟(格式YYYYMMDDHHMM)
              - description: 重新查询账号数据并推送诊断报告(python scripts/wechat_analyzer.py query --account_ids <redId>)
           b. 日程创建成功后,调用同步接口:
              ```bash
              python scripts/wechat_analyzer.py sync_notes --red_ids "redId"
              ```
           c. 同步接口:`https://redfox.hk/story/api/gzh/data/queryWorkList`,参数 `redId` 从账号信息中获取
           d. 输出提示:「订阅成功!系统将在30分钟后自动向您推送"xxx"账号的诊断报告,请耐心等待」
        3. 用户输入编号"2"或回复"仍然执行分析":
           - 继续执行分析流程
           - 在报告开头提示:「该账号暂未获取到近7天作品」
           - 输出完整的分析报告(爆文能力、更新产能等模块显示为"暂无作品数据")
        4. 用户输入编号"3"或回复"暂不订阅",结束本次分析
      
      **订阅推送机制(智能体执行)**:
      - 使用calendar_create工具创建日程任务
      - 日程触发时间:当前时间 + 30分钟(格式YYYYMMDDHHMM)
      - 日程触发后**无论是否有作品数据都执行分析流程**:
        1. 重新查询账号数据:`python scripts/wechat_analyzer.py query --account_ids <redId>`
        2. **如果works为空**:
           - 在报告开头提示:「该账号已重新同步,但暂未获取到近7天作品数据」
           - 继续执行完整的分析流程,输出诊断报告
        3. **如果works有数据**:
           - 正常输出完整的诊断报告
      
      ### 步骤3:在对话中输出诊断报告
      基于脚本输出的数据,直接在对话中输出四维度诊断报告。
      - **格式要求**:严格按照本文件"评分输出格式"章节的5节模板输出(账号信息→综合评分→近7天作品→优化建议→行业对标)
      - **输出位置**:直接在对话中输出,不创建markdown文件
      
      ### 步骤4:展示相似账号
      相似账号已整合在诊断报告第5版块中,与报告一体输出(无需单独展示):
      - 从脚本返回数据的 `similar_accounts` 字段提取最多5个相似账号
      - 按表格格式展示(序号/账号名称/平均阅读)
      - 最后输出"回复序号可继续分析!"
      
      - 该字段值为空未填充
      
      ### 步骤5:用户回复序号继续分析相似账号
      当用户回复序号选择某个相似账号时:
      1. 从 `similar_accounts` 数据中获取该账号的 `userId`
      2. 调用脚本查询该账号详情:
         ```bash
         python scripts/wechat_analyzer.py query --account_ids "userId"
         ```
      3. **执行完整分析流程**(同步骤3-5):
         - 步骤3:在对话中输出完整诊断报告
         - 步骤4:直接展示相似账号
         - 步骤5:生成HTML报告并展示
      4. **注意**:必须在对话中输出完整诊断报告,不能只输出HTML
      
      ## 评分体系
      
      ## 🎯 四维度诊断体系概览
      
      ### 综合评分 = 内容健康度×分类权重 + 用户活跃度×分类权重 + 内容核心数据×分类权重 + 运营规范性×分类权重(v4.1)
      
      ### 维度一:内容健康度(综合评分权重 25%)
      
      **评估账号内容生产能力和质量稳定性**
      
      | 评估项 | 权重 | 评分标准(0-10分) |
      |--------|------|-------------------|
      | 更新稳定性 | 15% | 10分=日更,8分=周更3-5次,5分=周更1-2次,<5=不规律 |
      | 内容垂直度 | 15% | 10分=极度垂直单一领域,7分=主领域+偶尔跨界,<5=内容杂乱 |
      | 原创能力 | 10% | 10分=90%+原创,8分=70-90%原创,5分=50-70%,<5=大量转载 |
      | 质量稳定性 | 10% | 阅读量波动系数(排除10万+截断值后计算),波动≤35%得高分;全部10万+直接满分 |
      | 内容深度 | 5% | 长文比例、专业引用、独家观点 |
      | 形式创新 | 5% | 多媒体运用、互动形式、排版创新 |
      
      **计算公式:**
      ```
      内容健康度 = 更新稳定性×0.15 + 内容垂直度×0.15 + 原创能力×0.10 + 质量稳定性×0.10 + 内容深度×0.05 + 形式创新×0.05
      ```
      
      ---
      
      ### 维度二:用户活跃度(综合评分权重 20%)
      
      **评估粉丝粘性和互动质量**
      
      | 评估项 | 权重 | 评分标准(0-10分) |
      |--------|------|-------------------|
      | 互动率 | 20% | (点赞+在看+留言)/阅读量:>5%得高分,<1%低分 |
      | 留言质量 | 10% | 评论密度(公众号评论需审核):≥0.05%=10分、≥0.02%=7、≥0.005%=4 |
      | 分享传播力 | 10% | 分享率:≥2%=10分、≥0.5%=7、≥0.2%=4(公众号2%+即为顶尖) |
      | 阅读完成率 | 5% | 基于文章长度和行业基准,>60%高分 |
      | 活跃时段集中度 | 5% | 推送时间固定性,用户阅读时间规律 |
      
      **计算公式:**
      ```
      用户活跃度 = 互动率×0.20 + 留言质量×0.10 + 分享传播力×0.10 + 阅读完成率×0.05 + 活跃时段×0.05
      ```
      
      > **v4.1 阅读量保底(分类感知)**:高均阅账号不应因低互动被过度惩罚。保底值按分类调整:时政号5万→≥70、3万→≥55;财经号5万→≥55、3万→≥45;通用5万→≥50、3万→≥40、1万→≥30。
      
      ---
      
      ## 综合评分体系(100分制)
      
      **综合评分体系:四维度加权均值**
      
      综合评分 = 内容健康度百分制×w1 + 用户活跃度百分制×w2 + 内容核心数据百分制×w3 + 运营规范性百分制×w4
      - v4.1: 权重按分类自适应(general=25/20/40/15, news_politics=25/8/52/15, finance_opinion=25/15/45/15, emotion_lifestyle=25/25/35/15, knowledge_edu=30/20/35/15, entertainment=20/15/45/20)
      - 各维度原始分转百分制:内容健康度(0-10)/10×100、用户活跃度(0-10)/10×100、核心数据(0-31)/31×100、运营规范性(0-10)/10×100
      
      | 评估模块 | 原始分 | 作用 | 对综合评分的影响 |
      |---------|--------|------|----------------|
      | **内容核心数据**(不含红狐) | 0-31分 | 调整分输入 | 微调 |
      | **内容健康度** | 0-10分 | 调整分输入 | 微调 |
      | **用户活跃度** | 0-10分 | 调整分输入 | 微调 |
      | **运营规范性** | 0-10分 | 调整分输入 | 微调 |
      
      ---
      
      ### 1. 内容健康度(0-10分原始分)
      
      **作为综合评分调整分输入之一**
      
      | 评估项 | 原始权重 | 评分标准(0-10分) |
      |--------|---------|-------------------|
      | 更新稳定性 | 15% | 10分=日更,8分=周更3-5次,5分=周更1-2次,<5=不规律 |
      | 内容垂直度 | 15% | 10分=极度垂直单一领域,7分=主领域+偶尔跨界,<5=内容杂乱 |
      | 原创能力 | 10% | 10分=90%+原创,8分=70-90%原创,5分=50-70%,<5=大量转载 |
      | 质量稳定性 | 10% | 阅读量波动系数(排除10万+截断值后计算),波动≤35%得高分;全部10万+直接满分 |
      | 内容深度 | 5% | 长文比例、专业引用、独家观点 |
      | 形式创新 | 5% | 多媒体运用、互动形式、排版创新 |
      
      **计算公式:** 内容健康度 = 更新稳定性x0.15 + 内容垂直度x0.15 + 原创能力x0.10 + 质量稳定性x0.10 + 内容深度x0.05 + 形式创新x0.05
      
      ---
      
      ### 2. 用户活跃度(0-10分原始分)
      
      **作为综合评分调整分输入之一**
      
      | 评估项 | 原始权重 | 评分标准(0-10分) |
      |--------|---------|-------------------|
      | 互动率 | 20% | (点赞+在看+留言)/阅读量:>5%得高分,<1%低分 |
      | 留言质量 | 10% | 评论密度(公众号评论需审核):≥0.05%=10分、≥0.02%=7、≥0.005%=4 |
      | 分享传播力 | 10% | 分享率:≥2%=10分、≥0.5%=7、≥0.2%=4(公众号2%+即为顶尖) |
      | 阅读完成率 | 5% | >60%高分,基于文章长度和行业基准 |
      | 活跃时段集中度 | 5% | 推送时间固定性,用户阅读时间规律 |
      
      **计算公式:** 用户活跃度 = 互动率x0.20 + 留言质量x0.10 + 分享传播力x0.10 + 阅读完成率x0.05 + 活跃时段x0.05
      
      ---
      
      ### 3. 内容核心数据表现
      
      **原始分0-31分,换算为百分制得分**
      
      | 评估项 | 分值 | 占比 | 说明 |
      |--------|------|------|------|
      | 爆款产出力 | 8分 | 25.8% | 10万+阅读文章占比 |
      | 阅读数表现 | 8分 | 25.8% | 内容传播广度 |
      | 点赞数表现 | 6分 | 19.4% | 内容认可度 |
      | 评论数表现 | 4分 | 12.9% | 用户参与度(微信评论需审核) |
      | 互动率表现 | 3分 | 9.7% | 综合互动质量 |
      | 发布时间合理性 | 2分 | 6.5% | 发布时机优化 |
      
      ---
      
      ### 4. 运营规范性(10分)
      
      | 评估项 | 分值 | 说明 |
      |--------|------|------|
      | 更新频率 | 5分 | 7天发文>=5篇=5分,3-4篇=4分,2篇=3分,1篇=2分 |
      | 发布时间合理性 | 3分 | 固定时段发文>=60%=3分,>=40%=2分,>=20%=1分 |
      | 账号认证 | 2分 | 有认证=2分,无=0分 |
      
      ---
      
      ### 综合评分计算
      
      综合评分 = 内容健康度百分制×w1 + 用户活跃度百分制×w2 + 内容核心数据百分制×w3 + 运营规范性百分制×w4
      - v4.1: 权重按分类自适应(general=25/20/40/15, news_politics=25/8/52/15, finance_opinion=25/15/45/15, emotion_lifestyle=25/25/35/15, knowledge_edu=30/20/35/15, entertainment=20/15/45/20)
      
      ---
      
      ### 综合评级标准(100分制)
      
      | 得分区间 | 评级 | 等级说明 | 特征描述 |
      |---------|------|---------|---------|
      | **80-100分** | 🏆 **标杆账号** | S级 | 均阅3万+、稳定日更、高互动的顶尖账号 |
      | **70-79分** | ⭐ **优质账号** | A级 | 均阅1-3万、持续优质内容的优质账号 |
      | **60-69分** | ✅ **健康账号** | B级 | 正常运营、有稳定输出的健康账号 |
      | **50-59分** | 📊 **中等账号** | C级 | 基础运营待提升 |
      | **40-49分** | ⚠️ **亚健康账号** | D级 | 运营不稳定或互动偏低,需调整改进 |
      | **<40分** | ❌ **问题账号** | E级 | 多项指标不达标,急需全面诊断 |
      
      ---
      
      ### 评分输出格式
      
      **严格按照以下5个章节顺序输出,不可省略任何章节。所有数值从脚本输出的JSON中对应字段取值,禁止自行编造。**
      
      #### JSON字段映射表
      
      脚本输出的JSON结构为 `{header, scores, content_health, user_activity, works, similar_accounts}`,报告各章节取值如下:
      
      | 报告位置 | 取值路径 | 说明 |
      |---------|---------|------|
      | 账号名称 | header.账号名 | |
      | 公众号ID | header.账号ID | |
      | 运营主体 | header.认证信息 | |
      | 账号简介 | header.账号简介 | |
      | 综合评分 | scores.综合评分 | 保留1位小数 |
      | 综合评级 | scores.综合评级 | 如"标杆账号" |
      | 综合等级 | scores.综合等级 | 如"S级" |
      | 近7天作品 | works | 数组,每项含标题(含超链接)/阅读数/点赞数/评论数/在看数/发布时间 |
      | 近7天作品提示 | works_hint | 无作品时为提示文字,否则为空
      | 行业对标 | scores.行业对标 | 含综合评分/平均阅读量/互动率/更新频率,每项有本账号/行业均值/头部账号 |
      | 相似账号 | similar_accounts | 数组,每项含账号名称/平均阅读数 |
      
      #### 一、账号信息(键值对,不使用表格)
      
      - **账号名称:** {header.账号名}
      - **公众号ID:** {header.账号ID}
      - **运营主体:** {header.认证信息}
      - **账号简介:** {header.账号简介}
      - **行业标签:** {header.账号标识}
      
      #### 二、综合评分
      
      **总分:{scores.综合评分} / 100分** — {scores.综合评级}({scores.综合等级})
      
      **行业对标表格:** 从scores.行业对标读取各指标的3个值(本账号/行业均值/头部账号),差距分析由智能体根据数据对比撰写。
      
      | 指标 | 本账号 | 行业均值 | 头部账号 | 差距分析 |
      |------|--------|---------|---------|---------|
      | **综合评分** | {scores.行业对标.综合评分.本账号} | {scores.行业对标.综合评分.行业均值} | {scores.行业对标.综合评分.头部账号} | [对比差距说明] |
      | **平均阅读量** | {scores.行业对标.平均阅读量.本账号} | {scores.行业对标.平均阅读量.行业均值} | {scores.行业对标.平均阅读量.头部账号} | [对比差距说明] |
      | **互动率** | {scores.行业对标.互动率.本账号} | {scores.行业对标.互动率.行业均值} | {scores.行业对标.互动率.头部账号} | [对比差距说明] |
      | **更新频率** | {scores.行业对标.更新频率.本账号} | {scores.行业对标.更新频率.行业均值} | {scores.行业对标.更新频率.头部账号} | [对比差距说明] |
      
      #### 三、近7天作品数据(最近5条)
      
      从works数组读取,每条作品输出6列。标题字段已包含Markdown超链接格式,直接原样输出。阅读数超过100000显示为"10w+",不足5条按实际条数输出,禁止增加其他列。
      
      > 若works_hint字段不为空,则不输出表格,仅输出works_hint中的提示文字。
      
      | 标题 | 阅读数 | 点赞数 | 评论数 | 在看数 | 发布时间 |
      |------|--------|--------|--------|--------|---------|
      | [示例文章标题](https://mp.weixin.qq.com/s/xxx) | 10w+ | 1280 | 356 | 89 | 05-22 08:30 |
      
      #### 四、优化建议
      
      分3个级别输出,基于各维度评分数据生成具体可执行建议:
      
      **紧急优化(1-2项)**
      1. **[具体问题]** → [可执行建议](预期:[量化效果])
      
      **重点优化(2-3项)**
      1. **[具体问题]** → [可执行建议](预期:[量化效果])
      
      **持续优化(1-2项)**
      1. **[优化方向]** → [策略建议](预期:[长期收益])
      
      #### 五、行业对标分析
      
      **对标结论:**
      - **综合评级:{scores.综合评级图标} {scores.综合等级}({scores.综合评分}分)**,[行业地位描述]
      - **核心优势:** [基于评分数据分析列出2-3个核心优势]
      - **与头部差距:** [说明与头部账号的差距]
      - **优化方向:** [基于评分数据分析列出2-3个优化方向]
      
      **相似账号:** 从similar_accounts数组读取,最多5条。
      
      | 序号 | 账号名称 | 平均阅读 |
      |------|---------|---------|---------|
      | 1 | [similar_accounts[0].账号名称] | [similar_accounts[0].平均阅读数] |
      | ... | ... | ... | ... |
      
      ---
      
      *诊断时间:[YYYY-MM-DD] | 数据来源:红狐API接口 | 置信度:[高/中] | 阅读数>10万显示为10w+ | 禁止输出西瓜指数*
      
      ## ⚠️ 注意事项
      
      ### 数据来源限制(重要)
      - 🔒 **接口数据唯一性:** 所有账号信息(基础信息、运营数据等)和内容文章数据(阅读数、点赞数、评论数等)**必须通过API接口获取**
      - ❌ **禁止自主采集:** 严禁从其他渠道(如新榜、清博、微信后台、第三方平台等)自主获取或估算数据
      - ✅ **接口优先原则:** 所有诊断分析必须基于接口返回的真实数据,不得使用估算值或推测值
      - 📊 **数据准确性:** 接口返回的数据为唯一可信数据源,所有评分和诊断结论必须基于此
      
      ### 数据准确性
      - 📊 **接口数据:** 所有数据来自红狐API接口,确保权威性和准确性
      - ⏰ **时效性:** 标注数据采集时间,提醒用户数据可能随时间变化
      - 🔍 **置信度标注:** 在报告末尾标注数据置信度(高=接口实时数据,中=接口缓存数据)
      - ⚠️ **未查询到账号:** 当接口返回空结果时,明确告知用户"未查询到该公众号信息",并提示"请输入准确的公众号名称或公众号ID重新查询",**严禁生成估算报告**
      
      ### 诊断客观性
      - ✅ **优势+问题并重:** 既要指出问题,也要肯定优势,避免片面评价
      - 📝 **数据支撑:** 所有诊断结论需有数据或现象支撑,不可主观臆断
      - 🎯 **可操作建议:** 建议必须具体可执行,包含预期效果,避免空泛表述
      - ⚖️ **评分一致性:** 同类型账号评分标准保持一致,确保横向可比
      
      ### 特殊场景处理
      - ❌ **未查询到账号:** 当无法找到用户输入的公众号信息,或**没有查询到公众号ID**时,明确告知用户"未查询到该公众号信息",并提示"请输入准确的公众号名称或公众号ID重新查询",**严禁生成估算报告**
      - 🔍 **新账号(<6个月):** 数据样本不足,降低诊断置信度至"中"或"低",侧重内容方向评估
      - 📉 **停更账号:** 标注"账号已停更",基于历史数据分析,提示数据时效性
      - 🎭 **矩阵账号:** 如属MCN矩阵,说明团队协作特点,评估需考虑资源协同效应
      - 💼 **企业号vs个人号:** 企业号侧重品牌传播和客户服务,个人号侧重IP打造和内容质量
      - 📰 **媒体号vs自媒体:** 媒体号有采编团队,标准更高;自媒体侧重个人特色和用户粘性
      
      ### 输出规范
      - 输出综合评分总分(100分制),不输出评分维度表格、优势模块和待优化模块
      - 🚫 **禁止输出内容健康度明细分析**,不得出现"内容健康度分析"独立章节或子项评分表格
      - 🚫 **禁止输出用户活跃度明细分析**,不得出现"用户活跃度分析"独立章节或子项评分表格
      - 📊 评分需精确到小数点后1位(如7.3/10)
      - 💡 优化建议按优先级分3级(紧急/重点/持续),数量分别为1-2项、2-3项、1-2项
      - 📈 行业对标需包含4个核心指标:综合评分、平均阅读量、互动率、更新频率
      - 📝 对标结论需包含:综合评级、核心优势、与头部差距、优化方向
      - 🎨 严格遵循Markdown格式,确保表格对齐、标题层级正确
      - 🔒 **所有数据必须来自接口**,不得使用估算值或外部数据源
      - 🚫 **禁止输出西瓜指数**,报告中不得出现任何与西瓜指数相关的字段或数据
      - 📖 **阅读数显示规则:** 阅读数超过100000时显示为"10w+"而非具体数字
      
      ### 常见行业基准参考(2024-2026)
      | 账号类型 | 平均阅读量 | 打开率 | 互动率 | 更新频率 |
      |---------|-----------|--------|--------|---------|
      | 头部大号 | 5万+ | 3-5% | 5-8% | 日更 |
      | 中腰部账号 | 5000-3万 | 2-4% | 3-5% | 周更3-5次 |
      | 小账号 | 500-3000 | 1-3% | 2-4% | 周更1-3次 |
      | B端专业号 | 2000-1万 | 5-10% | 2-3% | 周更2-4次 |
      | C端娱乐号 | 3000-5万 | 2-5% | 5-10% | 日更或隔日更 |
      
      ---
      
      ## 📚 使用示例
      
      ### 示例1:完整诊断(标准场景)
      ```
      用户:诊断"十点读书"公众号
      AI:[输出完整诊断报告,包含基本信息、核心数据、四维度诊断、优化建议、行业对标]
      ```
      
      ### 示例2:快速评估(简化场景)
      ```
      用户:分析一下"咪蒙"这个号
      AI:[输出完整诊断报告,可适当简化数据表格,但必须包含四维度诊断和优化建议]
      ```
      
      ### 示例3:对比诊断(多账号场景)
      ```
      用户:对比诊断"十点读书"和"洞见"
      AI:[分别输出两个账号的完整诊断报告] + [横向对比总结表格] + [差异化建议]
      ```
      
      ### 示例4:使用公众号ID查询
      ```
      用户:诊断公众号ID为"isdushu"的账号
      AI:[先确认账号名称,再输出完整诊断报告]
      ```
      
      ### 示例5:特殊场景处理
      ```
      用户:帮我看看"XX公众号"怎么样
      AI:[检测到新账号/停更账号等特殊情况,在报告中标注置信度和特殊说明]
      ```
      
      ### 示例6:未查询到账号
      ```
      用户:诊断"不存在的公众号123"
      AI:抱歉,未查询到"不存在的公众号123"的公众号信息。请输入准确的公众号名称或公众号ID重新查询。
      ```
      
  • scripts
    • analyzer.py 24.9 KB
      """analyzer.py — 公众号账号分析编排
      
      职责:单账号评分处理、原始数据保存、查询命令、同步命令。
      被 wechat_analyzer.py(入口)调用。
      """
      
      import json
      import os
      
      from api_client import (
          https_post,
          API_PATH_SEARCH_USER,
          API_PATH_QUERY_DATA,
          RAW_DATA_FILE,
      )
      from scoring import (
          _score_content_health,
          _score_user_activity,
          _score_core_data,
          _score_operation_compliance,
          _get_score_level,
          _get_score_level_icon,
          _get_overall_grade,
          _work_read,
          _work_like,
          _work_comment,
          _work_share,
          _work_publish_time,
          _calc_avg_read,
          _classify_account_category,
          CATEGORY_WEIGHTS,
          CATEGORY_ACTIVITY_FLOOR,
          _ACTIVITY_FLOOR_GENERAL,
          CATEGORY_NAMES,
      )
      from report import REPORT_DATA_FILE
      
      
      # ════════════════════════════════════════════════════════════
      #  单账号分析
      # ════════════════════════════════════════════════════════════
      
      def _analyze_single_account(raw, has_works=True):
          """对单个账号原始数据进行评分和结构化处理
      
          适配 queryData 接口返回的字段命名:
          accountId / accountName / avatar / verifyName / description / works / similarAccounts
          """
          # 无作品提示
          no_works_hint = "" if has_works else "该账号暂未获取到作品数据"
      
          # 统一字段读取
          nickname = raw.get("accountName", "")
          account_id = raw.get("accountId", "")
          signature = raw.get("description", "")
          avg_read_count = raw.get("avgReadCount", 0) or 0
          account_type = raw.get("accountType", "")
          verify_name = raw.get("verifyName", "")
          works = raw.get("works", []) or []
      
          if not works:
              has_works = False
      
          # v4.1: 账号分类识别(在评分之前调用)
          category_info = _classify_account_category(account_type, signature, works, verify_name)
          category = category_info["category"]
          category_name = category_info["category_name"]
      
          # 四维度评分(用户活跃度传入分类参数)
          content_health = _score_content_health(works, signature, verify_name, account_type)
          user_activity = _score_user_activity(works if has_works else [], category=category)
          core_data = _score_core_data(works if has_works else [], avg_read_count, category=category)
          operation_compliance = _score_operation_compliance(works if has_works else [], verify_name)
      
          # 综合评分:v4.1 分类自适应权重
          weights = CATEGORY_WEIGHTS.get(category, CATEGORY_WEIGHTS["general"])
          dim1_score = round(content_health["原始分"] / 10 * 100, 1)
          dim2_score = round(user_activity["原始分"] / 10 * 100, 1)
          dim3_score = round(core_data["原始分"] / 31 * 100, 1)
          dim4_score = round(operation_compliance["原始分"] / 10 * 100, 1)
      
          # v4.1: 分类感知的阅读量保底机制
          # 时政号保底更激进(互动天然极低),财经号次之,其他分类使用 general 保底
          floor_config = CATEGORY_ACTIVITY_FLOOR.get(category, _ACTIVITY_FLOOR_GENERAL)
          for threshold in sorted(floor_config.keys(), reverse=True):
              if avg_read_count >= threshold:
                  dim2_score = max(dim2_score, floor_config[threshold])
                  break
      
          total_score = round(
              dim1_score * weights["content_health"]
              + dim2_score * weights["user_activity"]
              + dim3_score * weights["core_data"]
              + dim4_score * weights["operation_compliance"], 1
          )
          total_score = max(0, min(100, total_score))
      
          # 评级
          overall_level = _get_score_level(total_score, 100)
          overall_grade, overall_rank = _get_overall_grade(total_score)
          dim1_level = _get_score_level(dim1_score, 100)
          dim1_level_icon = _get_score_level_icon(dim1_score, 100)
          dim2_level = _get_score_level(dim2_score, 100)
          dim2_level_icon = _get_score_level_icon(dim2_score, 100)
          dim3_level = _get_score_level(dim3_score, 100)
          dim3_level_icon = _get_score_level_icon(dim3_score, 100)
          dim4_level = _get_score_level(dim4_score, 100)
          dim4_level_icon = _get_score_level_icon(dim4_score, 100)
      
          # 构建scores结构
          scores = {
              "综合评分": total_score,
              "综合得分层级": overall_level,
              "综合评级": overall_grade,
              "综合等级": overall_rank,
              "账号分类": category_name,
              "分类ID": category,
              "使用权重": {
                  "内容健康度": f"{int(weights['content_health']*100)}%",
                  "用户活跃度": f"{int(weights['user_activity']*100)}%",
                  "内容核心数据": f"{int(weights['core_data']*100)}%",
                  "运营规范性": f"{int(weights['operation_compliance']*100)}%",
              },
              "内容健康度得分": dim1_score,
              "内容健康度满分": 100,
              "内容健康度得分率": round(dim1_score, 1),
              "内容健康度评级": dim1_level,
              "内容健康度评级图标": dim1_level_icon,
              "用户活跃度得分": dim2_score,
              "用户活跃度满分": 100,
              "用户活跃度得分率": round(dim2_score, 1),
              "用户活跃度评级": dim2_level,
              "用户活跃度评级图标": dim2_level_icon,
              "内容核心数据表现得分": dim3_score,
              "内容核心数据表现满分": 100,
              "内容核心数据表现得分率": round(dim3_score, 1),
              "内容核心数据表现评级": dim3_level,
              "内容核心数据表现评级图标": dim3_level_icon,
              "运营规范性得分": dim4_score,
              "运营规范性满分": 100,
              "运营规范性得分率": round(dim4_score, 1),
              "运营规范性评级": dim4_level,
              "运营规范性评级图标": dim4_level_icon,
              # 内容健康度子项(原始0-10分,×3=0-30分)
              "更新稳定性得分": content_health["更新稳定性"],
              "更新稳定性满分": 10,
              "内容垂直度得分": content_health["内容垂直度"],
              "内容垂直度满分": 10,
              "原创能力得分": content_health["原创能力"],
              "原创能力满分": 10,
              "质量稳定性得分": content_health["质量稳定性"],
              "质量稳定性满分": 10,
              "内容深度得分": content_health["内容深度"],
              "内容深度满分": 10,
              "形式创新得分": content_health["形式创新"],
              "形式创新满分": 10,
              "内容健康度原始分": content_health["原始分"],
              # 用户活跃度子项(原始0-10分,×2.5=0-25分)
              "互动率得分": user_activity["互动率"],
              "互动率满分": 10,
              "留言质量得分": user_activity["留言质量"],
              "留言质量满分": 10,
              "分享传播力得分": user_activity["分享传播力"],
              "分享传播力满分": 10,
              "阅读完成率得分": user_activity["阅读完成率"],
              "阅读完成率满分": 10,
              "活跃时段集中度得分": user_activity["活跃时段集中度"],
              "活跃时段集中度满分": 10,
              "用户活跃度原始分": user_activity["原始分"],
              # 内容核心数据表现子项(0-43分制)
              "阅读数表现得分": core_data["阅读数表现"],
              "阅读数表现满分": 8,
              "点赞数表现得分": core_data["点赞数表现"],
              "点赞数表现满分": 6,
              "评论数表现得分": core_data["评论数表现"],
              "评论数表现满分": 4,
              "互动率表现得分": core_data["互动率表现"],
              "互动率表现满分": 3,
              "发布时间合理性得分": core_data["发布时间合理性"],
              "发布时间合理性满分": 2,
              "爆款产出力得分": core_data["爆款产出力"],
              "爆款产出力满分": 8,
              "爆款率": core_data.get("爆款率", 0),
              "内容核心数据原始分": core_data["原始分"],
              # 运营规范性子项(直接0-10分)
              "更新频率得分": operation_compliance["更新频率"],
              "更新频率满分": 5,
              "发布时间合理性2得分": operation_compliance["发布时间合理性"],
              "发布时间合理性2满分": 3,
              "账号认证得分": operation_compliance["账号认证"],
              "账号认证满分": 2,
          }
      
          # 计算优势模块和待优化模块
          dim_score_rates = [
              {"维度名": "内容健康度", "得分": dim1_score, "得分率": dim1_score},
              {"维度名": "用户活跃度", "得分": dim2_score, "得分率": dim2_score},
              {"维度名": "内容核心数据表现", "得分": dim3_score, "得分率": dim3_score},
              {"维度名": "运营规范性", "得分": dim4_score, "得分率": dim4_score},
          ]
          dim_sorted = sorted(dim_score_rates, key=lambda x: x["得分率"], reverse=True)
          scores["优势模块"] = dim_sorted[:2]
          scores["待优化模块"] = dim_sorted[-2:]
      
          # 互动率和更新频率计算
          # 注意:queryData 的 interactiveCount = 阅读数+互动分项之和,不能直接用于互动率计算
          # 始终使用分项加总:like+comment+share+watch
          # 阅读数含10万+截断值(100001为保守下界),与 _score_user_activity 口径保持一致
          interaction_rate = 0
          if works:
              total_reads = sum(_work_read(w) for w in works if _work_read(w) > 0)
              total_interactions = sum(
                  _work_like(w) + _work_comment(w) + _work_share(w) + (w.get("watchCount") or 0)
                  for w in works
              )
              if total_reads > 0:
                  interaction_rate = round(total_interactions / total_reads * 100, 2)
          works_7d = len(works) if works else 0
      
          # 行业对标(基于公众号真实行业水平)
          # 数据来源:公众号行业研究报告,覆盖10万+个人/企业公众号
          scores["行业对标"] = {
              "综合评分": {"本账号": f"{total_score}分", "行业均值": "45-55分", "头部账号": "85-95分"},
              "平均阅读量": {"本账号": str(avg_read_count), "行业均值": "2000-8000", "头部账号": "3-10万"},
              "互动率": {"本账号": f"{interaction_rate}%", "行业均值": "1-3%", "头部账号": "10-25%"},
              "更新频率": {"本账号": f"{works_7d}篇/近期", "行业均值": "2-4篇/周", "头部账号": "5-7篇/周"},
          }
      
          # 构建返回结果
          result = {
              "header": {
                  "账号名": nickname,
                  "账号ID": account_id,
                  "账号链接": f"https://open.weixin.qq.com/qr/code?username={account_id}" if account_id else "",
                  "账号类型": account_type,
                  "账号分类": category_name,
                  "分类ID": category,
                  "分类置信度": category_info["confidence"],
                  "账号简介": signature,
                  "认证信息": verify_name,
                  "平均阅读数": avg_read_count,
                  "no_works_hint": no_works_hint,
              },
              "scores": scores,
              "content_health": {
                  "更新稳定性": f"{content_health['更新稳定性']}/10",
                  "内容垂直度": f"{content_health['内容垂直度']}/10",
                  "原创能力": f"{content_health['原创能力']}/10",
                  "质量稳定性": f"{content_health['质量稳定性']}/10",
                  "内容深度": f"{content_health['内容深度']}/10",
                  "形式创新": f"{content_health['形式创新']}/10",
                  "原始分": f"{content_health['原始分']}/10",
                  "总分": f"{dim1_score}/50",
                  "评级": dim1_level,
                  "评级图标": dim1_level_icon,
              },
              "user_activity": {
                  "互动率": f"{user_activity['互动率']}/10",
                  "留言质量": f"{user_activity['留言质量']}/10",
                  "分享传播力": f"{user_activity['分享传播力']}/10",
                  "阅读完成率": f"{user_activity['阅读完成率']}/10",
                  "活跃时段集中度": f"{user_activity['活跃时段集中度']}/10",
                  "原始分": f"{user_activity['原始分']}/10",
                  "总分": f"{dim2_score}/50",
                  "评级": dim2_level,
                  "评级图标": dim2_level_icon,
              },
              "works": [
                  {
                      "标题": ("[{0}]({1})".format(w.get("title", ""), w.get("workUrl", "")) if w.get("workUrl") else w.get("title", "")),
                      "阅读数": _work_read(w),
                      "点赞数": _work_like(w),
                      "评论数": _work_comment(w),
                      "在看数": w.get("watchCount") or 0,
                      "发布时间": _work_publish_time(w),
                  }
                  for w in (works[:5] if works else [])
              ],
              "works_hint": "",
              "similar_accounts": [
                  {
                      "账号名称": sa.get("accountName", ""),
                      "账号ID": sa.get("accountId", ""),
                      "平均阅读数": "",  # queryData 相似账号不返回该字段
                  }
                  for sa in ((raw.get("similarAccounts") or [])[:5])
                  if sa.get("accountName")
              ],
              "_raw": raw,
          }
      
          # 保存report_data.json
          script_dir = os.path.dirname(os.path.abspath(__file__))
          output_dir = os.path.normpath(os.path.join(script_dir, "..", "output"))
          os.makedirs(output_dir, exist_ok=True)
          report_path = os.path.join(output_dir, REPORT_DATA_FILE)
          with open(report_path, "w", encoding="utf-8") as f:
              json.dump(result, f, ensure_ascii=False, indent=2)
      
          return result
      
      
      def _save_raw_data(raw_data):
          """将接口原始数据保存到raw_data.json"""
          script_dir = os.path.dirname(os.path.abspath(__file__))
          output_dir = os.path.join(script_dir, "..", "output")
          output_dir = os.path.normpath(output_dir)
          os.makedirs(output_dir, exist_ok=True)
          raw_path = os.path.join(output_dir, RAW_DATA_FILE)
          with open(raw_path, "w", encoding="utf-8") as f:
              json.dump(raw_data, f, ensure_ascii=False, indent=2)
      
      
      # ════════════════════════════════════════════════════════════
      #  查询命令
      # ════════════════════════════════════════════════════════════
      
      def cmd_query(account_ids=None, account_names=None, force_analyze=False):
          """查询命令:searchUser 获取微信号 → queryData 精确查询完整数据
      
          流程:
          1. searchUser(keyword=name) 搜索名称 → 必须名称完全匹配,取 account(微信号)
          2. queryData(accountIds=[微信号], accountNames=[名称]) → 精确定位,返回完整数据+works+相似账号
      
          按 ID 查询时:预期传入微信号(如 duhaoshu)或 gh_xxx 格式
      
          Args:
              account_ids: 微信号/gh_xxx ID 列表(可选)
              account_names: 账号名称列表(可选)
              force_analyze: 是否强制分析
          """
          items = []
          names_to_query = account_names if account_names else []
          ids_to_query = account_ids if account_ids else []
      
          if not names_to_query and not ids_to_query:
              print(json.dumps({
                  "status": "error",
                  "message": "请提供账号名称或账号ID",
                  "query_type": "not_found",
                  "data": []
              }, ensure_ascii=False))
              return
      
          # ── 按名称查询:searchUser 微信号 → queryData 完整数据 ──
          for name in names_to_query:
              # 第一步:searchUser 搜索,必须名称完全匹配
              try:
                  search_resp = https_post(API_PATH_SEARCH_USER, {"keyword": name, "offset": 0})
              except Exception as e:
                  print(json.dumps({
                      "status": "error",
                      "message": f"搜索失败: {str(e)}",
                      "query_type": "not_found",
                      "data": []
                  }, ensure_ascii=False))
                  return
      
              if not (isinstance(search_resp, dict) and search_resp.get("code") == 2000):
                  print(json.dumps({
                      "status": "error",
                      "message": f"搜索失败: {search_resp.get('msg', '未知错误')}",
                      "query_type": "not_found",
                      "data": []
                  }, ensure_ascii=False))
                  return
      
              search_list = (search_resp.get("data") or {}).get("list", [])
              if not search_list:
                  print(json.dumps({
                      "status": "success",
                      "query_type": "not_found",
                      "message": f"未查询到账号【{name}】,该账号可能尚未收录或名称有误,请核实公众号名称后重试。",
                      "data": []
                  }, ensure_ascii=False))
                  return
      
              # 必须名称完全匹配
              matched = next((a for a in search_list if a.get("accountName", "") == name), None)
              if matched is None:
                  candidates = [
                      {
                          "name": a.get("accountName", ""),
                          "wxId": a.get("wxId", ""),       # 公众号ID,gh_xxx 格式
                          "account": a.get("account", ""),  # 微信号,如 duhaoshu
                      }
                      for a in search_list[:5]
                      if a.get("accountName")
                  ]
                  # 候选列表文案:账号名称 + 公众号ID + 微信号
                  candidates_lines = []
                  for c in candidates:
                      parts = [f"「{c['name']}」"]
                      if c["wxId"]:
                          parts.append(f"ID: {c['wxId']}")
                      if c["account"]:
                          parts.append(f"微信号: {c['account']}")
                      candidates_lines.append(" ".join(parts))
                  candidates_str = "、".join(candidates_lines)
                  print(json.dumps({
                      "status": "success",
                      "query_type": "not_found",
                      "message": (
                          f"未找到名称为「{name}」的公众号,"
                          f"搜索结果中有较相近的账号:{candidates_str},"
                          f"请确认公众号名称后重试,或可用公众号ID(微信号)直接查询。"
                      ),
                      "candidates": candidates,
                      "data": []
                  }, ensure_ascii=False))
                  return
      
              # 微信号(account 字段)= queryData.accountIds 的正确参数格式
              weixin_id = matched.get("account", "")
      
              # 第二步:queryData 用微信号+名称精确定位,得到 works+similarAccounts
              try:
                  data_resp = https_post(API_PATH_QUERY_DATA, {
                      "accountIds": [weixin_id] if weixin_id else [],
                      "accountNames": [name]
                  })
              except Exception as e:
                  print(json.dumps({
                      "status": "error",
                      "message": f"请求失败: {str(e)}",
                      "query_type": "not_found",
                      "data": []
                  }, ensure_ascii=False))
                  return
      
              if not (isinstance(data_resp, dict) and data_resp.get("code") == 2000):
                  print(json.dumps({
                      "status": "error",
                      "message": f"查询失败: {data_resp.get('msg', '未知错误')}",
                      "query_type": "not_found",
                      "data": []
                  }, ensure_ascii=False))
                  return
      
              data_list = data_resp.get("data") or []
              if not data_list:
                  print(json.dumps({
                      "status": "success",
                      "query_type": "not_found",
                      "message": f"未查询到账号【{name}】的详细数据。",
                      "data": []
                  }, ensure_ascii=False))
                  return
      
              # 优先取微信号匹配的账号,否则取 works 最多的
              exact = next((a for a in data_list if a.get("account", "") == weixin_id), None)
              best = exact if exact else max(data_list, key=lambda x: len(x.get("works", []) or []))
              works = best.get("works", []) or []
              avg_read = _calc_avg_read(works)
              items.append({**best, "avgReadCount": avg_read})
      
          # ── 按 ID 查询:直接用微信号/gh_xxx 调 queryData ──
          for wx_id in ids_to_query:
              try:
                  resp = https_post(API_PATH_QUERY_DATA, {
                      "accountIds": [wx_id],
                      "accountNames": []
                  })
              except Exception as e:
                  print(json.dumps({
                      "status": "error",
                      "message": f"请求失败: {str(e)}",
                      "query_type": "not_found",
                      "data": []
                  }, ensure_ascii=False))
                  return
      
              if not (isinstance(resp, dict) and resp.get("code") == 2000):
                  print(json.dumps({
                      "status": "error",
                      "message": f"查询失败: {resp.get('msg', '未知错误')}",
                      "query_type": "not_found",
                      "data": []
                  }, ensure_ascii=False))
                  return
      
              data_list = resp.get("data") or []
              if not data_list:
                  print(json.dumps({
                      "status": "success",
                      "query_type": "not_found",
                      "message": f"未查询到ID为【{wx_id}】的公众号,请确认ID是否正确。",
                      "data": []
                  }, ensure_ascii=False))
                  return
      
              account = data_list[0]
              works = account.get("works", []) or []
              avg_read = _calc_avg_read(works)
              items.append({**account, "avgReadCount": avg_read})
      
      
          if not items:
              print(json.dumps({
                  "status": "success",
                  "query_type": "not_found",
                  "data": []
              }, ensure_ascii=False))
              return
      
          _save_raw_data(items)
      
          accounts_with_works = [it for it in items if it.get("works")]
          accounts_need_sync = [it for it in items if not it.get("works")]
      
          if len(items) > 1:
              if not accounts_with_works:
                  if force_analyze:
                      [_analyze_single_account(it, has_works=False) for it in items]
                      print(json.dumps({
                          "status": "success",
                          "query_type": "multi",
                          "message": "数据已保存",
                          "no_works_hint": "暂未获取到作品数据"
                      }, ensure_ascii=False))
                      return
                  print(json.dumps({
                      "status": "success",
                      "query_type": "need_sync",
                      "message": "这些账号暂无作品数据",
                      "need_sync": [{"nickname": it.get("accountName", ""), "redId": it.get("accountId", "")} for it in items]
                  }, ensure_ascii=False))
                  return
      
              [_analyze_single_account(it) for it in accounts_with_works]
              output = {"status": "success", "query_type": "multi", "message": "数据已保存"}
              if accounts_need_sync:
                  output["need_sync"] = [{"nickname": it.get("accountName", ""), "redId": it.get("accountId", "")} for it in accounts_need_sync]
              print(json.dumps(output, ensure_ascii=False))
              return
      
          raw_item = items[0]
          works = raw_item.get("works", []) or []
          _analyze_single_account(raw_item, has_works=bool(works))
          print(json.dumps({
              "status": "success",
              "query_type": "single",
              "message": "数据已保存",
              "no_works_hint": "该账号暂无作品数据" if not works else None
          }, ensure_ascii=False))
      
      
      # ════════════════════════════════════════════════════════════
      #  同步命令
      # ════════════════════════════════════════════════════════════
      
      def cmd_sync_notes(account_ids):
          """订阅命令:调用接口同步账号作品数据
      
          参数:
              account_ids: 公众号账号ID列表
          """
          results = []
      
          for account_id in account_ids:
              try:
                  body = {
                      "accountId": account_id,
                      "source": "公众号账号诊断-GitHub"
                  }
                  response = https_post("/story/api/gzhUser/syncUserNotes", body)
      
                  if isinstance(response, dict) and response.get("code") == 5000:
                      results.append({
                          "accountId": account_id,
                          "account_name": f"账号{account_id}",
                          "status": "success",
                          "schedule_required": True,
                          "schedule_time_minutes": 30
                      })
                  else:
                      results.append({
                          "accountId": account_id,
                          "account_name": f"账号{account_id}",
                          "status": "success",
                          "schedule_required": True,
                          "schedule_time_minutes": 30
                      })
              except Exception as e:
                  results.append({
                      "accountId": account_id,
                      "status": "error",
                      "message": f"订阅失败: {str(e)}"
                  })
      
          print(json.dumps({
              "status": "success",
              "query_type": "sync",
              "data": {"sync_results": results}
          }, ensure_ascii=False))
      
    • api_client.py 2.5 KB
      """api_client.py — 红狐 API HTTP 通信层
      
      职责:API 凭证管理、HTTP POST 请求封装。
      被 analyzer.py 和 report.py 调用。
      """
      
      import json
      import os
      import re
      import sys
      
      import requests
      
      
      # ── API 常量 ──
      API_HOST = "redfox.hk"
      API_PATH_SEARCH_USER = "/story/api/gzh/data/searchUser"  # 接口1:关键词搜索账号 → 获取微信号
      API_PATH_QUERY_DATA = "/story/api/gzhUser/queryData"      # 接口2:按微信号+名称精确查询完整数据
      RAW_DATA_FILE = "raw_data.json"
      
      
      def _read_from_shell_config():
          """从shell配置文件中尝试读取REDFOX_API_KEY(仅macOS/Linux)"""
          if sys.platform == "win32":
              return None
          home = os.path.expanduser("~")
          config_files = [
              os.path.join(home, ".zshrc"),
              os.path.join(home, ".bashrc"),
              os.path.join(home, ".bash_profile"),
              os.path.join(home, ".profile"),
          ]
          for config_file in config_files:
              try:
                  if os.path.isfile(config_file):
                      with open(config_file, "r", encoding="utf-8", errors="ignore") as f:
                          content = f.read()
                      match = re.search(r'export\s+REDFOX_API_KEY\s*=\s*["\']?([^"\'\n]+)["\']?', content)
                      if match:
                          return match.group(1).strip()
              except (OSError, PermissionError):
                  continue
          return None
      
      
      def _get_credential():
          """获取API凭证 - 优先从环境变量REDFOX_API_KEY读取,其次从shell配置文件读取"""
          credential = os.getenv("REDFOX_API_KEY")
          if credential and credential.strip():
              return credential.strip()
      
          # 环境变量未设置,尝试从shell配置文件读取
          credential = _read_from_shell_config()
          if credential:
              return credential
      
          raise ValueError(
              "未找到 REDFOX_API_KEY,请配置环境变量后重试。\n"
              "  macOS/Linux: export REDFOX_API_KEY=<你的apikey>\n"
              "  Windows:     [Environment]::SetEnvironmentVariable('REDFOX_API_KEY', '<值>', 'User')\n"
              "获取API Key: 访问 https://redfox.hk/ 注册后在个人中心获取"
          )
      
      
      def _get_headers():
          """获取请求头"""
          return {
              "Content-Type": "application/json",
              "Accept": "application/json",
              "X-API-KEY": _get_credential(),
          }
      
      
      def https_post(path, body_dict):
          """POST请求"""
          url = f"https://{API_HOST}{path}"
          body_json = json.dumps(body_dict, ensure_ascii=False)
          response = requests.post(url, data=body_json.encode("utf-8"), headers=_get_headers(), timeout=30)
          return response.json()
      
    • report.py 26.4 KB
      """report.py — HTML 报告生成
      
      职责:模板替换、条件区域控制、单账号/多账号 HTML 报告生成。
      被 wechat_analyzer.py(入口)和 analyzer.py 调用。
      """
      
      import json
      import os
      import re
      import sys
      from datetime import datetime
      
      from scoring import _format_interactive_count
      
      
      # ── 常量 ──
      REPORT_DATA_FILE = "report_data.json"
      MULTI_REPORT_DATA_FILE = "multi_report_data.json"
      
      # 各维度满分映射
      SCORE_MAX_MAP = {
          "内容健康度": 100,
          "用户活跃度": 100,
          "内容核心数据表现": 100,
          "运营规范性": 100,
      }
      
      
      # ════════════════════════════════════════════════════════════
      #  HTML 模板辅助函数
      # ════════════════════════════════════════════════════════════
      
      def _flatten_dict(d, parent_key="", sep="."):
          """将嵌套字典扁平化为点分隔的键名"""
          items = []
          for k, v in d.items():
              new_key = f"{parent_key}{sep}{k}" if parent_key else k
              if isinstance(v, dict):
                  items.extend(_flatten_dict(v, new_key, sep).items())
              elif isinstance(v, list):
                  for i, item in enumerate(v):
                      if isinstance(item, dict):
                          items.extend(_flatten_dict(item, f"{new_key}[{i}]", sep).items())
                      else:
                          items.append((f"{new_key}[{i}]", str(item) if item is not None else ""))
              else:
                  items.append((new_key, str(v) if v is not None else ""))
          return dict(items)
      
      
      def _build_replacements(report_data):
          """构建HTML模板替换字典"""
          replacements = {}
      
          # 近期作品表格行
          works_rows = []
          for w in report_data.get("works", []):
              if not isinstance(w, dict):
                  continue
              title = w.get("title", "无标题")[:15].replace(" ", "")
              if not title.strip():
                  title = "无标题"
              date_str = w.get("date", "-") or "-"
              likes = w.get("likes", "0") or "0"
              url = w.get("workUrl", "") or ""
              link = f'<a href="{url}" target="_blank">查看</a>' if url else "-"
              works_rows.append(f"<tr><td>{title}</td><td>{date_str}</td><td>{likes}</td><td>{link}</td></tr>")
          replacements["{{works_table_rows}}"] = "\n".join(works_rows)
      
          # 爆文列表表格
          viral_list = report_data.get("viral", {}).get("爆文列表", [])
          if not viral_list:
              viral_list = report_data.get("爆文列表", [])
          viral_rows = []
          if viral_list:
              for v in viral_list:
                  if not isinstance(v, dict):
                      continue
                  title = v.get("标题", v.get("title", "-"))[:20] or "-"
                  pub_time = v.get("发布时间", v.get("publishTime", "-")) or "-"
                  interactive = v.get("互动数", v.get("interactiveCount", "-")) or "-"
                  multiple = v.get("超标准倍数", v.get("multiple", "-")) or "-"
                  viral_rows.append(f"<tr><td>{title}</td><td>{pub_time}</td><td>{interactive}</td><td>{multiple}</td></tr>")
          if viral_rows:
              viral_table_html = (
                  '<table class="viral-table">\n'
                  '    <tr><th>爆文标题</th><th>发布时间</th><th>互动数</th><th>超标准倍数</th></tr>\n'
                  + "\n".join(viral_rows) + "\n"
                  '</table>'
              )
          else:
              viral_table_html = '<div class="info-row"><span class="label">爆文列表:</span><span class="value">暂无爆文</span></div>'
          replacements["{{爆文列表表格}}"] = viral_table_html
      
          return replacements
      
      
      def _is_empty_field(val):
          """判断字段是否为空"""
          return str(val).strip() in ("", "None", "none")
      
      
      def _remove_section_markers(html, marker_name, should_show):
          """根据条件移除或保留标记区域"""
          start_tag = f"<!-- {marker_name}_START -->"
          end_tag = f"<!-- {marker_name}_END -->"
          if should_show:
              html = html.replace(start_tag, "").replace(end_tag, "")
          else:
              html = re.sub(rf'<!-- {marker_name}_START -->.*?<!-- {marker_name}_END -->', '', html, flags=re.DOTALL).rstrip()
          return html
      
      
      def _remove_empty_info_rows(html):
          """移除HTML中值为空的info-row行"""
          html = re.sub(r'<div class="info-row">\s*<span class="label">[^<]*</span>\s*<span class="value">\s*</span>\s*</div>', '', html)
          return html
      
      
      def _remove_conditional_sections(html, report_data):
          """根据数据条件移除空数据模块"""
          scores = report_data.get("scores", {})
      
          # 爆文能力:始终展示,移除条件隐藏
          # viral_count = scores.get("爆文数", "")
          # try:
          #     viral_val = int(viral_count) if not _is_empty_field(viral_count) else 0
          # except (ValueError, TypeError):
          #     viral_val = 0
          # html = _remove_section_markers(html, "SECTION_VIRAL", viral_val > 0)
          # 直接移除标记,始终显示爆文能力模块
          html = html.replace("<!-- SECTION_VIRAL_START -->", "").replace("<!-- SECTION_VIRAL_END -->", "")
      
          # 近期作品:works为空时隐藏
          works = report_data.get("works", [])
          has_valid_works = any(isinstance(w, dict) and w.get("title", "").strip() for w in works)
          html = _remove_section_markers(html, "SECTION_WORKS", has_valid_works)
      
          # 可强化:内容健康度<16分时显示
          account_health = report_data.get("content_health", {})
          scores = report_data.get("scores", {})
          health_score = scores.get("内容健康度得分", 0)
          show_can_enhance = health_score < 16
          html = _remove_section_markers(html, "SECTION_CAN_ENHANCE", show_can_enhance)
      
          return html
      
      
      # ════════════════════════════════════════════════════════════
      #  单账号 HTML 报告生成
      # ════════════════════════════════════════════════════════════
      
      def cmd_generate_html():
          """生成单账号HTML命令 - 直接使用report_data.json中的similar_accounts数据"""
          script_dir = os.path.dirname(os.path.abspath(__file__))
          data_path = os.path.normpath(os.path.join(script_dir, "..", "output", REPORT_DATA_FILE))
          template_path = os.path.normpath(os.path.join(script_dir, "..", "assets", "report_template.html"))
          raw_data_path = os.path.normpath(os.path.join(script_dir, "..", "output", "raw_data.json"))
      
          if not os.path.exists(data_path):
              print(json.dumps({"status": "error", "message": f"报告数据文件不存在: {data_path},请先完成诊断报告生成并保存report_data.json"}, ensure_ascii=False))
              sys.exit(1)
      
          if not os.path.exists(template_path):
              print(json.dumps({"status": "error", "message": f"模板文件不存在: {template_path}"}, ensure_ascii=False))
              sys.exit(1)
      
          with open(data_path, "r", encoding="utf-8") as f:
              report_data = json.load(f)
      
          # 读取原始数据作为备用数据源
          raw_data = {}
          if os.path.exists(raw_data_path):
              try:
                  with open(raw_data_path, "r", encoding="utf-8") as f:
                      raw_data = json.load(f)
              except Exception:
                  pass  # 原始数据读取失败时忽略
      
          with open(template_path, "r", encoding="utf-8") as f:
              html = f.read()
      
          replacements = _build_replacements(report_data)
      
          # 相似账号卡片(账号名称为超链接)- 直接使用report_data.json中的similar_accounts
          # 每个账号一行,展示:账号名称(超链接)、平均阅读、总互动、推荐理由、发文特点、可学之处
          similar_cards = []
          for sa in report_data.get("similar_accounts", []):
              if not isinstance(sa, dict):
                  continue
              name = sa.get("账号名称") or sa.get("accountName") or sa.get("nickname") or sa.get("name", "")
              account_url = sa.get("账号链接") or sa.get("profileUrl") or sa.get("url", "")
              avg_read = sa.get("平均阅读数") or sa.get("avgReadCount") or 0
              total_interactive = sa.get("总互动") or sa.get("totalInteractiveCount") or sa.get("interactiveCountThirty", 0)
              recommend_reason = sa.get("推荐理由") or sa.get("recommendReason") or ""
              post_feature = sa.get("发文特点") or sa.get("postFeature") or ""
              learn_point = sa.get("可学之处") or sa.get("learnPoint") or ""
      
              # 账号名称超链接
              name_html = f'<a href="{account_url}" target="_blank" style="color:#1890ff;text-decoration:none;font-weight:500;">{name}</a>' if account_url else name
              # 如果没有链接但有accountId,构造链接
              if not account_url:
                  account_id = sa.get("accountId") or sa.get("redId") or sa.get("userId", "")
                  if account_id:
                      account_url = f"https://mp.weixin.qq.com/profile/{account_id}"
                      name_html = f'<a href="{account_url}" target="_blank" style="color:#1890ff;text-decoration:none;font-weight:500;">{name}</a>'
      
              similar_cards.append(
                  f'<div class="similar-card-row" style="padding:12px 0;border-bottom:1px solid #f0f0f0;">'
                  f'<div style="margin-bottom:6px;"><strong>{name_html}</strong> | 平均阅读:{avg_read} | 总互动:{total_interactive}</div>'
                  f'<div style="font-size:13px;color:#666;margin-bottom:4px;"><strong>推荐理由:</strong>{recommend_reason}</div>'
                  f'<div style="font-size:13px;color:#666;margin-bottom:4px;"><strong>发文特点:</strong>{post_feature}</div>'
                  f'<div style="font-size:13px;color:#666;"><strong>可学之处:</strong>{learn_point}</div>'
                  f'</div>'
              )
          replacements["{{similar_accounts_cards}}"] = "\n".join(similar_cards)
      
          # 执行替换
          for key, val in replacements.items():
              html = html.replace(key, val)
      
          # 条件移除空数据模块
          html = _remove_conditional_sections(html, report_data)
      
          # 相似账号区域 - 直接展示(移除条件注释)
          html = html.replace("<!-- SIMILAR_START -->", "").replace("<!-- SIMILAR_END -->", "")
      
          html = _remove_empty_info_rows(html)
      
          # ========== 自检:检测未替换的模板字段 ==========
          unreplaced = re.findall(r'\{\{[^}]+\}\}', html)
          if unreplaced:
              # 收集所有未替换的字段
              unique_unreplaced = list(set(unreplaced))
              # 扁平化分析数据和原始数据
              flat_data = _flatten_dict(report_data)
              _rd = raw_data[0] if isinstance(raw_data, list) and len(raw_data) > 0 else raw_data
              flat_raw = _flatten_dict(_rd) if _rd and isinstance(_rd, dict) else {}
      
              # 字段名映射:模板字段名 -> 原始数据字段名
              field_mapping = {
                  "总在看数": "collected",
                  "总点赞数": "liked",
                  "近30天互动量": "interactions_30d",
                  "近30天发作品数": "works_30d",
                  "作品总数": "works_total",
                  "账号名": "nickname",
                  "官方等级": "level",
              }
      
              for field in unique_unreplaced:
                  field_name = field[2:-2]  # 移除 {{ 和 }}
                  found_value = None
      
                  # 第一步:从分析数据中查找
                  for k, v in flat_data.items():
                      if k == field_name or k.endswith("." + field_name):
                          if v is not None and str(v).strip() != "" and str(v) != "0":
                              found_value = v
                              break
      
                  # 第二步:分析数据为空,从原始数据中查找
                  if found_value is None and flat_raw:
                      # 先尝试字段名映射
                      if field_name in field_mapping:
                          raw_field = field_mapping[field_name]
                          for k, v in flat_raw.items():
                              if k == raw_field or k.endswith("." + raw_field):
                                  if v is not None and str(v).strip() != "":
                                      found_value = v
                                      # 格式化大数字
                                      if field_name in ["总在看数", "总点赞数", "近30天互动量"]:
                                          found_value = _format_interactive_count(v)
                                      break
                      # 再尝试直接匹配字段名
                      if found_value is None:
                          for k, v in flat_raw.items():
                              if k == field_name or k.endswith("." + field_name):
                                  if v is not None and str(v).strip() != "":
                                      found_value = v
                                      break
      
                  # 第三步:根据值进行处理
                  if found_value is not None and str(found_value).strip() != "":
                      html = html.replace(field, str(found_value))
                  else:
                      # 数据中无值,根据字段类型填充默认值
                      numeric_fields = ["得分", "分", "互动", "收藏", "点赞", "数", "率", "量", "倍", "篇", "天", "中位数参考", "优秀值参考", "等级"]
                      is_numeric = any(nf in field_name for nf in numeric_fields)
                      if is_numeric:
                          html = html.replace(field, "0")
                      else:
                          html = html.replace(field, "")
      
              # 再次检查是否还有未替换字段
              remaining = re.findall(r'\{\{[^}]+\}\}', html)
              if remaining:
                  print(json.dumps({
                      "status": "error",
                      "message": f"HTML模板字段未完全替换: {list(set(remaining))}",
                      "unreplaced_fields": list(set(remaining))
                  }, ensure_ascii=False))
                  sys.exit(1)
      
          # 输出HTML文件
          output_dir = os.path.normpath(os.path.join(script_dir, "..", "output"))
          os.makedirs(output_dir, exist_ok=True)
      
          account_name = report_data.get("header", {}).get("账号名", "report")
          safe_name = account_name.replace("/", "_").replace("\\", "_").replace(" ", "_")
          timestamp = datetime.now().strftime("%Y%m%d_%H%M%S")
          output_path = os.path.join(output_dir, f"{safe_name}_诊断报告_{timestamp}.html")
          with open(output_path, "w", encoding="utf-8") as f:
              f.write(html)
      
          result_info = {
              "status": "success",
              "message": "HTML报告已生成",
              "output_path": output_path
          }
          print(json.dumps(result_info, ensure_ascii=False))
      
      
      # ════════════════════════════════════════════════════════════
      #  多账号对比 HTML 报告生成
      # ════════════════════════════════════════════════════════════
      
      def _build_account_detail_html(account_data, account_index):
          """为多账号报告生成单个账号的详情HTML"""
          replacements = _build_replacements(account_data)
      
          # 评分条
          for dim, max_score in SCORE_MAX_MAP.items():
              score_key = dim + "得分"
              pct_key = dim + "得分_pct"
              score_val = replacements.get("{{" + score_key + "}}", "0")
              try:
                  pct = round(int(score_val) / max_score * 100)
              except (ValueError, ZeroDivisionError):
                  pct = 0
              replacements["{{" + pct_key + "}}"] = str(pct)
      
          header = account_data.get("header", {})
          raw_data = account_data.get("_raw", {})
          avatar = raw_data.get("头像", "")
          name = header.get("账号名", "")
          tag = header.get("账号标识", "")
          score = replacements.get("{{综合评分}}", "-")
      
          # 构建详情HTML(简化版,用于多账号对比)
          detail = f'''  <div class="account-detail">
          <div class="account-detail-header">
            <img src="{avatar}" onerror="this.style.display='none'">
            <span class="name">{name}</span>
            <span class="tag">{tag}</span>
            <span class="score">{score}分</span>
          </div>
          <div class="account-detail-body">
            <div class="score-bars">
              <div class="score-bar-item"><span class="name">内容健康度</span><div class="bar-bg"><div class="bar-fill" style="width:{replacements.get("{{内容健康度得分_pct}}", "0")}%"></div></div><span class="val">{replacements.get("{{内容健康度得分}}", "0")}分</span></div>
              <div class="score-bar-item"><span class="name">用户活跃度</span><div class="bar-bg"><div class="bar-fill" style="width:{replacements.get("{{用户活跃度得分_pct}}", "0")}%"></div></div><span class="val">{replacements.get("{{用户活跃度得分}}", "0")}分</span></div>
      
            </div>
      
            <div style="margin-top:12px; font-weight:600; font-size:13px; color:#FF2442;">综合诊断</div>
            <div class="conclusion">
              <p>{replacements.get("{{综合诊断结论内容}}", "")}</p>
            </div>'''
      
          # 行动处方
          detail += f'''
            <div style="margin-top:10px; font-weight:600; font-size:13px; color:#FF2442;">行动处方</div>
            <div class="action-item"><strong>问题归因</strong>:<br>• {replacements.get("{{问题归因1}}", "")}<br>• {replacements.get("{{问题归因2}}", "")}</div>
            <div class="action-item"><strong>具体动作</strong>:<br>1. {replacements.get("{{具体动作1}}", "")}<br>2. {replacements.get("{{具体动作2}}", "")}<br>3. {replacements.get("{{具体动作3}}", "")}</div>'''
      
          detail += '''
          </div>
        </div>'''
      
          detail = _remove_empty_info_rows(detail)
          return detail
      
      
      def cmd_generate_multi_html(with_similar=False):
          """生成多账号对比HTML命令"""
          script_dir = os.path.dirname(os.path.abspath(__file__))
          data_path = os.path.normpath(os.path.join(script_dir, "..", "output", MULTI_REPORT_DATA_FILE))
          template_path = os.path.normpath(os.path.join(script_dir, "..", "assets", "multi_report_template.html"))
      
          if not os.path.exists(data_path):
              print(json.dumps({"status": "error", "message": f"多账号报告数据文件不存在: {data_path},请先保存multi_report_data.json"}, ensure_ascii=False))
              sys.exit(1)
      
          if not os.path.exists(template_path):
              print(json.dumps({"status": "error", "message": f"多账号模板文件不存在: {template_path}"}, ensure_ascii=False))
              sys.exit(1)
      
          with open(data_path, "r", encoding="utf-8") as f:
              multi_data = json.load(f)
      
          with open(template_path, "r", encoding="utf-8") as f:
              html = f.read()
      
          accounts = multi_data.get("accounts", [])
          if not accounts:
              print(json.dumps({"status": "error", "message": "accounts数组为空,无账号数据"}, ensure_ascii=False))
              sys.exit(1)
      
          html = html.replace("{{账号数量}}", str(len(accounts)))
      
          data_time = multi_data.get("header", {}).get("数据获取时间", "")
          if not data_time and accounts:
              data_time = accounts[0].get("header", {}).get("数据获取时间", "")
          html = html.replace("{{数据获取时间}}", data_time)
      
          # 对比表头
          header_cells = []
          for acc in accounts:
              name = acc.get("header", {}).get("账号名", "")
              header_cells.append(f"<th>{name}</th>")
          html = html.replace("{{对比表头}}", "".join(header_cells))
      
          # 对比表格行
          compare_rows = []
          metrics = [
              ("综合评分", "scores", "综合评分"),
              ("平均阅读数", "data_performance", "平均阅读数"),
              ("互动率", "user_activity", "互动率"),
              ("内容健康度", "scores", "内容健康度得分"),
              "separator",
              ("用户活跃度", "scores", "用户活跃度得分"),
          ]
      
          best_map = {}
          for item in metrics:
              if item == "separator":
                  continue
              label, section, key = item
              values = []
              for acc in accounts:
                  sec = acc.get(section, {})
                  raw_val = sec.get(key, "")
                  try:
                      v = float(raw_val) if str(raw_val).strip() not in ("",) else None
                  except (ValueError, TypeError):
                      v = None
                  values.append(v)
              valid_vals = [v for v in values if v is not None]
              if valid_vals:
                  best_map[key] = max(valid_vals)
      
          for item in metrics:
              if item == "separator":
                  compare_rows.append('<tr style="height:4px;background:#FDE8EC;"><td colspan="99"></td></tr>')
                  continue
              label, section, key = item
              cells = [f"<td>{label}</td>"]
              for acc in accounts:
                  sec = acc.get(section, {})
                  raw_val = sec.get(key, "")
                  val_str = str(raw_val) if raw_val else ""
                  if val_str.strip() == "":
                      cells.append("<td>-</td>")
                      continue
                  try:
                      num_val = float(raw_val) if str(raw_val).strip() not in ("",) else None
                  except (ValueError, TypeError):
                      num_val = None
                  if num_val is not None and key in best_map and num_val == best_map[key]:
                      cells.append(f'<td class="best">{val_str}</td>')
                  else:
                      cells.append(f"<td>{val_str}</td>")
              compare_rows.append("<tr>" + "".join(cells) + "</tr>")
          html = html.replace("{{对比表格行}}", "\n".join(compare_rows))
      
          # 对比总结
          comparison = multi_data.get("comparison", {})
      
          diff_items = comparison.get("核心差异", [])
          diff_html = ""
          if diff_items:
              diff_html = '<div class="summary-module summary-diff"><div class="module-title"><span class="icon">⚡</span> 核心差异</div>'
              for item in diff_items:
                  if isinstance(item, dict):
                      acc_name = item.get("账号名", "")
                      content = item.get("内容", "")
                      if content:
                          diff_html += f'<div style="margin-bottom:8px; padding:6px 10px; background:#fff; border-radius:6px;"><span style="font-weight:600; color:#D48806; font-size:12px;">{acc_name}</span><p style="font-size:13px; color:#555; line-height:1.8; margin:4px 0 0;">{content}</p></div>'
              diff_html += '</div>'
          html = html.replace("{{对比总结_核心差异}}", diff_html)
      
          common_items = comparison.get("共同问题", [])
          common_html = ""
          if common_items:
              common_html = '<div class="summary-module summary-common"><div class="module-title"><span class="icon">🔗</span> 共同问题</div><ul style="padding-left:18px; margin:0;">'
              for item in common_items:
                  if isinstance(item, str) and item.strip():
                      common_html += f'<li style="font-size:13px; color:#555; line-height:1.8;">{item}</li>'
              common_html += '</ul></div>'
          html = html.replace("{{对比总结_共同问题}}", common_html)
      
          advice_items = comparison.get("发展建议", [])
          advice_html = ""
          if advice_items:
              advice_html = '<div class="summary-module summary-advice"><div class="module-title"><span class="icon">🚀</span> 发展建议</div>'
              for item in advice_items:
                  if isinstance(item, dict):
                      acc_name = item.get("账号名", "")
                      content = item.get("内容", "")
                      if content:
                          advice_html += f'<div style="margin-bottom:8px; padding:6px 10px; background:#fff; border-radius:6px;"><span style="font-weight:600; color:#FF2442; font-size:12px;">{acc_name}</span><p style="font-size:13px; color:#555; line-height:1.8; margin:4px 0 0;">{content}</p></div>'
              advice_html += '</div>'
          html = html.replace("{{对比总结_发展建议}}", advice_html)
      
          # 各账号详情
          details = []
          for i, acc in enumerate(accounts):
              details.append(_build_account_detail_html(acc, i))
          html = html.replace("{{各账号详情}}", "\n".join(details))
      
          # 条件移除空数据模块
          html = _remove_conditional_sections(html, multi_data)
          html = _remove_empty_info_rows(html)
      
          # ========== 自检:检测未替换的模板字段 ==========
          unreplaced = re.findall(r'\{\{[^}]+\}\}', html)
          if unreplaced:
              unique_unreplaced = list(set(unreplaced))
              # 对于空值字段,根据数据类型处理
              flat_multi = _flatten_dict(multi_data)
              for field in unique_unreplaced:
                  field_name = field[2:-2]  # 移除 {{ 和 }}
                  # 在扁平化数据中查找对应值
                  found_value = None
                  for k, v in flat_multi.items():
                      if k == field_name or k.endswith("." + field_name):
                          found_value = v
                          break
                  # 也在各账号数据中查找
                  if found_value is None or found_value == "":
                      for acc in accounts:
                          flat_acc = _flatten_dict(acc)
                          for k, v in flat_acc.items():
                              if k == field_name or k.endswith("." + field_name):
                                  found_value = v
                                  break
                          if found_value is not None and found_value != "":
                              break
                  if found_value is not None and found_value != "":
                      # 数据中有值,使用该值
                      html = html.replace(field, str(found_value))
                  else:
                      # 数据中无值,根据字段类型填充默认值
                      numeric_fields = ["得分", "分", "互动", "收藏", "点赞", "数", "率", "量", "倍", "篇", "天", "中位数参考", "优秀值参考", "等级"]
                      is_numeric = any(nf in field_name for nf in numeric_fields)
                      if is_numeric:
                          html = html.replace(field, "0")
                      else:
                          html = html.replace(field, "")
              remaining = re.findall(r'\{\{[^}]+\}\}', html)
              if remaining:
                  print(json.dumps({
                      "status": "error",
                      "message": f"HTML模板字段未完全替换: {list(set(remaining))}",
                      "unreplaced_fields": list(set(remaining))
                  }, ensure_ascii=False))
                  sys.exit(1)
      
          # 输出HTML文件
          output_dir = os.path.normpath(os.path.join(script_dir, "..", "output"))
          os.makedirs(output_dir, exist_ok=True)
      
          names = [acc.get("header", {}).get("账号名", "未知") for acc in accounts[:3]]
          safe_name = "vs".join(n.replace("/", "_").replace("\\", "_").replace(" ", "_") for n in names)
          timestamp = datetime.now().strftime("%Y%m%d_%H%M%S")
          output_path = os.path.join(output_dir, f"{safe_name}_对比报告_{timestamp}.html")
          with open(output_path, "w", encoding="utf-8") as f:
              f.write(html)
      
          result_info = {
              "status": "success",
              "message": "多账号对比HTML报告已生成",
              "output_path": output_path
          }
          print(json.dumps(result_info, ensure_ascii=False))
      
    • scoring.py 39.6 KB
      """scoring.py — 公众号账号评分引擎
      
      职责:作品数据辅助函数、四维度评分、等级判定、格式化工具。
      被 analyzer.py 和 report.py 调用。
      """
      
      from datetime import datetime
      
      
      # ════════════════════════════════════════════════════════════
      #  v4.1 分类自适应评分配置
      # ════════════════════════════════════════════════════════════
      
      # 分类关键词库(description + 作品标题匹配)
      CATEGORY_KEYWORDS = {
          "news_politics": [
              "时政", "新闻", "决策", "政策", "国际", "局势", "天下", "时事", "政治",
              "党", "政府", "两会", "战略", "策赢", "决胜", "环球", "观察", "人民日报",
          ],
          "finance_opinion": [
              "财经", "投资", "经济", "金融", "商业", "股市", "理财", "资本", "市场",
              "A股", "基金", "证券", "宏观经济", "吴晓波",
          ],
          "emotion_lifestyle": [
              "情感", "读书", "生活", "美文", "故事", "深夜", "陪伴", "观点", "洞见",
              "成长", "心灵", "温暖", "治愈", "励志", "十点",
          ],
          "knowledge_edu": [
              "知识", "教育", "学习", "科普", "书店", "文化", "历史", "学术", "研究",
              "阅读", "人文",
          ],
          "entertainment": [
              "搞笑", "娱乐", "段子", "视频", "明星", "八卦", "综艺", "电影", "追剧",
          ],
      }
      
      # 分类名称(中文映射,用于报告输出)
      CATEGORY_NAMES = {
          "news_politics": "时政新闻",
          "finance_opinion": "财经理政",
          "emotion_lifestyle": "情感生活",
          "knowledge_edu": "知识教育",
          "entertainment": "娱乐休闲",
          "general": "综合",
      }
      
      # 各分类的四维度权重(总和必须=1.00)
      CATEGORY_WEIGHTS = {
          "general":           {"content_health": 0.25, "user_activity": 0.20, "core_data": 0.40, "operation_compliance": 0.15},
          "news_politics":     {"content_health": 0.25, "user_activity": 0.08, "core_data": 0.52, "operation_compliance": 0.15},
          "finance_opinion":   {"content_health": 0.25, "user_activity": 0.15, "core_data": 0.45, "operation_compliance": 0.15},
          "emotion_lifestyle": {"content_health": 0.25, "user_activity": 0.25, "core_data": 0.35, "operation_compliance": 0.15},
          "knowledge_edu":     {"content_health": 0.30, "user_activity": 0.20, "core_data": 0.35, "operation_compliance": 0.15},
          "entertainment":     {"content_health": 0.20, "user_activity": 0.15, "core_data": 0.45, "operation_compliance": 0.20},
      }
      
      # 各分类的用户活跃度保底值(均阅阈值 → 百分制下限)
      CATEGORY_ACTIVITY_FLOOR = {
          "news_politics":     {50000: 70.0, 30000: 55.0, 10000: 40.0},
          "finance_opinion":   {50000: 55.0, 30000: 45.0, 10000: 30.0},
          # 其他分类使用 general 保底
      }
      
      # general 保底(默认)
      _ACTIVITY_FLOOR_GENERAL = {50000: 50.0, 30000: 40.0, 10000: 30.0}
      
      # 各分类的用户活跃度子项阈值(互动率、评论密度、分享率)
      # 格式:{score_value: threshold},从高到低匹配
      CATEGORY_ACTIVITY_THRESHOLDS = {
          "general": {
              "interaction": {1.0: 0.05, 0.7: 0.02, 0.4: 0.01},
              "comment":     {1.0: 0.0005, 0.7: 0.0002, 0.4: 0.00005},
              "share":       {1.0: 0.02, 0.7: 0.005, 0.4: 0.002},
          },
          "news_politics": {
              "interaction": {1.0: 0.03, 0.7: 0.012, 0.4: 0.006},
              "comment":     {1.0: 0.0002, 0.7: 0.00008, 0.4: 0.00002},
              "share":       {1.0: 0.015, 0.7: 0.003, 0.4: 0.001},
          },
          "finance_opinion": {
              "interaction": {1.0: 0.04, 0.7: 0.015, 0.4: 0.008},
              "comment":     {1.0: 0.0003, 0.7: 0.0001, 0.4: 0.00003},
              "share":       {1.0: 0.015, 0.7: 0.004, 0.4: 0.0015},
          },
          "emotion_lifestyle": {
              "interaction": {1.0: 0.05, 0.7: 0.02, 0.4: 0.01},
              "comment":     {1.0: 0.0005, 0.7: 0.0002, 0.4: 0.00005},
              "share":       {1.0: 0.02, 0.7: 0.005, 0.4: 0.002},
          },
          "knowledge_edu": {
              "interaction": {1.0: 0.04, 0.7: 0.02, 0.4: 0.01},
              "comment":     {1.0: 0.0004, 0.7: 0.00015, 0.4: 0.00004},
              "share":       {1.0: 0.02, 0.7: 0.005, 0.4: 0.002},
          },
          "entertainment": {
              "interaction": {1.0: 0.06, 0.7: 0.03, 0.4: 0.015},
              "comment":     {1.0: 0.0005, 0.7: 0.0002, 0.4: 0.00005},
              "share":       {1.0: 0.03, 0.7: 0.01, 0.4: 0.003},
          },
      }
      
      
      def _classify_account_category(account_type, signature, works, verify_name=""):
          """v4.1: 根据账号信息自动识别内容分类
      
          三级信号融合:accountType → description关键词 → 作品标题关键词 + 互动结构验证
          返回:{"category": str, "confidence": float, "category_name": str, "signals": dict}
          """
          signals = {"account_type": account_type, "desc_match": "", "title_match": "", "behavior_match": ""}
          scores_by_cat = {}
      
          # 第一级:accountType 字段直接映射(若非空)
          _type_map = {
              "新闻媒体": "news_politics", "时政": "news_politics", "新闻": "news_politics",
              "财经": "finance_opinion", "金融": "finance_opinion",
              "情感": "emotion_lifestyle", "生活": "emotion_lifestyle",
              "教育": "knowledge_edu", "文化": "knowledge_edu",
              "娱乐": "entertainment", "影视": "entertainment",
          }
          if account_type and account_type.strip():
              for key, cat in _type_map.items():
                  if key in account_type:
                      return {"category": cat, "confidence": 0.95, "category_name": CATEGORY_NAMES.get(cat, "综合"),
                              "signals": {**signals, "account_type": account_type}}
      
          # 第二级:description 关键词匹配
          desc = (signature or "").lower()
          desc_scores = {}
          desc_match_parts = []
          for cat, keywords in CATEGORY_KEYWORDS.items():
              hits = [kw for kw in keywords if kw.lower() in desc]
              if hits:
                  desc_scores[cat] = len(hits)
                  desc_match_parts.append(f"{cat}:[{','.join(hits[:3])}]")
          if desc_match_parts:
              signals["desc_match"] = "; ".join(desc_match_parts)
      
          # 第三级:作品标题关键词统计
          title_scores = {}
          title_match_parts = []
          if works:
              titles = [w.get("title", "") for w in works[:10] if w.get("title")]
              for cat, keywords in CATEGORY_KEYWORDS.items():
                  hit_count = sum(1 for t in titles if any(kw.lower() in t.lower() for kw in keywords))
                  if hit_count > 0:
                      title_scores[cat] = hit_count / max(len(titles), 1)
                      title_match_parts.append(f"{cat}:{hit_count}/{len(titles)}")
          if title_match_parts:
              signals["title_match"] = "; ".join(title_match_parts)
      
          # 第四级:互动结构验证
          behavior_scores = {}
          if works:
              total_reads = sum(w.get("clicksCount", 0) or 0 for w in works)
              total_comments = sum(w.get("commentCount", 0) or 0 for w in works)
              avg_read = total_reads / len(works) if works else 0
              comment_rate = total_comments / total_reads if total_reads > 0 else 0
              # 评论率极低 + 高阅读 → 时政号特征
              if avg_read >= 10000 and comment_rate < 0.0001:
                  behavior_scores["news_politics"] = 0.8
                  signals["behavior_match"] = f"高均阅({int(avg_read)})+极低评论率({comment_rate:.4%})"
              elif avg_read >= 10000 and comment_rate < 0.0003:
                  behavior_scores["finance_opinion"] = 0.6
                  signals["behavior_match"] = f"高均阅({int(avg_read)})+低评论率({comment_rate:.4%})"
              elif comment_rate >= 0.0003:
                  behavior_scores["emotion_lifestyle"] = 0.5
                  signals["behavior_match"] = f"评论率较高({comment_rate:.4%})"
      
          # 综合加权:description(0.5) + title(0.3) + behavior(0.2)
          # 使用 CATEGORY_KEYWORDS 的插入顺序作为等分 tie-break 优先级(确定性)
          all_cats_ordered = []
          for cat in list(CATEGORY_KEYWORDS.keys()):
              if cat in desc_scores or cat in title_scores or cat in behavior_scores:
                  if cat not in all_cats_ordered:
                      all_cats_ordered.append(cat)
          for cat in all_cats_ordered:
              d = desc_scores.get(cat, 0)
              # description命中数归一化(最多3个即满分)
              d_norm = min(d / 3, 1.0) if d > 0 else 0
              t_norm = title_scores.get(cat, 0)
              b_norm = behavior_scores.get(cat, 0)
              scores_by_cat[cat] = d_norm * 0.5 + min(t_norm, 1.0) * 0.3 + b_norm * 0.2
      
          if not scores_by_cat:
              return {"category": "general", "confidence": 1.0, "category_name": "综合", "signals": signals}
      
          # 等分时按 CATEGORY_KEYWORDS 插入顺序优先(确定性 tie-break)
          best_cat = max(all_cats_ordered, key=lambda c: scores_by_cat.get(c, 0))
          best_score = scores_by_cat[best_cat]
      
          # 置信度 >= 0.3 即确认分类(description有1个关键词命中即可)
          if best_score >= 0.3:
              return {"category": best_cat, "confidence": round(best_score, 2),
                      "category_name": CATEGORY_NAMES.get(best_cat, "综合"), "signals": signals}
      
          return {"category": "general", "confidence": 1.0, "category_name": "综合", "signals": signals}
      
      
      # ════════════════════════════════════════════════════════════
      #  作品数据辅助函数
      # ════════════════════════════════════════════════════════════
      
      def _work_read(w):
          """获取作品阅读数(兼容多种字段名)"""
          for key in ("clicksCount", "readCount", "readNum"):
              val = w.get(key)
              if val is not None:
                  return val
          return 0
      
      
      def _work_like(w):
          """获取作品点赞数"""
          for key in ("likeCount", "likedCount"):
              val = w.get(key)
              if val is not None:
                  return val
          return 0
      
      
      def _work_comment(w):
          """获取作品评论数"""
          return w.get("commentCount") or 0
      
      
      def _work_share(w):
          """获取作品分享数"""
          return w.get("shareCount") or 0
      
      
      def _work_collect(w):
          """获取作品在看数(微信以'在看'近似收藏行为)"""
          return w.get("watchCount") or 0
      
      
      def _work_interact_total(w):
          """获取作品总互动数 = 点赞+评论+分享+在看"""
          return _work_like(w) + _work_comment(w) + _work_share(w) + _work_collect(w)
      
      
      def _work_publish_time(w):
          """获取作品发布时间"""
          return w.get("publishTime") or w.get("time") or w.get("timestamp") or w.get("createTime") or ""
      
      
      # ════════════════════════════════════════════════════════════
      #  基准数据与交互结构
      # ════════════════════════════════════════════════════════════
      
      def _extract_benchmark_from_api(raw):
          """从接口数据中提取水平衡量基准数据
      
          Args:
              raw: 接口返回的原始数据,包含 accountAvgList 和 accountExcellentList
      
          Returns:
              dict: benchmark字典
          """
          avg_list = raw.get("accountAvgList", []) or []
          excellent_list = raw.get("accountExcellentList", []) or []
      
          # 构建中位数参考字典(取第一条数据)
          avg_dict = {}
          if avg_list:
              for k, v in avg_list[0].items():
                  if k != "fansType":
                      try:
                          avg_dict[k] = float(v) if v else 0
                      except (ValueError, TypeError):
                          pass
      
          # 构建优秀值参考字典(取第一条数据)
          excellent_dict = {}
          if excellent_list:
              for k, v in excellent_list[0].items():
                  if k != "fansType":
                      try:
                          excellent_dict[k] = float(v) if v else 0
                      except (ValueError, TypeError):
                          pass
      
          # 映射到水平衡量指标
          benchmark = {
              "近30天作品互动量": {
                  "中位数参考": avg_dict.get("近30天作品互动量均值", 0),
                  "优秀值参考": excellent_dict.get("近30天作品互动量均值", 0),
              },
              "近30天发作品数": {
                  "中位数参考": avg_dict.get("近30天发作品数均值", 0),
                  "优秀值参考": excellent_dict.get("近30天发作品数均值", 0),
              },
              "总点赞数": {
                  "中位数参考": avg_dict.get("总点赞数均值", 0),
                  "优秀值参考": excellent_dict.get("总点赞数均值", 0),
              },
              "总在看数": {
                  "中位数参考": avg_dict.get("总在看数均值", 0),
                  "优秀值参考": excellent_dict.get("总在看数均值", 0),
              },
              "作品总数": {
                  "中位数参考": avg_dict.get("作品总数均值", 0),
                  "优秀值参考": excellent_dict.get("作品总数均值", 0),
              },
              "周更频率": {
                  "中位数参考": avg_dict.get("近30天发作品数均值", 0) / 4.0 if avg_dict.get("近30天发作品数均值", 0) else 2.0,
                  "优秀值参考": excellent_dict.get("近30天发作品数均值", 0) / 4.0 if excellent_dict.get("近30天发作品数均值", 0) else 5.0,
              },
          }
      
          # 计算互动率和收藏率(基于阅读数)
          # 互动率 = (点赞数 + 在看数) / 阅读数 * 100%
          # 收藏率 = 在看数 / 阅读数 * 100%
      
          # 获取中位数参考的点赞数、在看数
          avg_like = avg_dict.get("总点赞数均值", 0)
          avg_collect = avg_dict.get("总在看数均值", 0)
          avg_read = avg_dict.get("平均阅读数均值", 100)
      
          # 获取优秀值参考的点赞数、在看数
          excellent_like = excellent_dict.get("总点赞数均值", 0)
          excellent_collect = excellent_dict.get("总在看数均值", 0)
          excellent_read = excellent_dict.get("平均阅读数均值", 1000)
      
          # 计算互动率中位数参考和优秀值参考(基于阅读数)
          interaction_rate_avg = round((avg_like + avg_collect) / avg_read * 100, 2) if avg_read > 0 else 0.5
          interaction_rate_excellent = round((excellent_like + excellent_collect) / excellent_read * 100, 2) if excellent_read > 0 else 1.5
      
          # 计算收藏率中位数参考和优秀值参考
          collect_rate_avg = round(avg_collect / avg_read * 100, 2) if avg_read > 0 else 1.0
          collect_rate_excellent = round(excellent_collect / excellent_read * 100, 2) if excellent_read > 0 else 3.0
      
          benchmark["互动率"] = {
              "中位数参考": interaction_rate_avg,
              "优秀值参考": interaction_rate_excellent,
          }
          benchmark["收藏率"] = {
              "中位数参考": collect_rate_avg,
              "优秀值参考": collect_rate_excellent,
          }
      
          return benchmark
      
      
      def _calc_interaction_structure(works):
          """计算互动结构(点赞/评论/在看占比)"""
          if not works:
              return None, None
      
          total_like = sum(_work_like(w) for w in works)
          total_comment = sum(_work_comment(w) for w in works)
          total_share = sum(_work_share(w) for w in works)
      
          total = total_like + total_comment + total_share
          if total == 0:
              return None, None
      
          like_pct = round(total_like / total * 100, 1)
          collect_pct = round((total_comment + total_share) / total * 100, 1)
      
          return like_pct, collect_pct
      
      
      def _get_level_judgment(value, benchmark, is_lower_better=False):
          """根据基准值判断等级
      
          基准数据结构:
          - 中位数参考:同层级账号中位数
          - 优秀值参考:同层级账号优秀值
          """
          if value is None:
              return "数据不足"
      
          median_val = benchmark.get("中位数参考", 0)
          excellent_val = benchmark.get("优秀值参考", 0)
      
          if is_lower_better:
              # 数值越低越好(如间隔标准差)
              if value <= excellent_val:
                  return "优秀"
              elif value <= median_val:
                  return "良好"
              else:
                  return "待提升"
          else:
              # 数值越高越好(如互动率、收藏率)
              if value >= excellent_val:
                  return "优秀"
              elif value >= median_val:
                  return "良好"
              else:
                  return "待提升"
      
      
      def _format_interactive_count(count):
          """互动量格式化,>=10000转w+,<10000直接展示原值"""
          try:
              count = int(count)
          except (ValueError, TypeError):
              return str(count)
          if count >= 10000:
              w_val = count / 10000
              if w_val == int(w_val):
                  return f"{int(w_val)}w+"
              return f"{round(w_val, 1)}w+"
          else:
              return str(count)
      
      
      def _calc_avg_read(works):
          """从作品列表计算平均阅读数
      
          10万+文章按100001计入均值:该值为微信阅读量显示上限的截断值,
          真实阅读量≥10万,属于保守下界。排除爆款文章会系统性低估头部账号
          (如爆款率25%的账号)的平均阅读水平。
          """
          if not works:
              return 0
          reads = [_work_read(w) for w in works if _work_read(w) > 0]
          return int(sum(reads) / len(reads)) if reads else 0
      
      
      def _calc_viral_ratio(works):
          """计算爆款率:10万+阅读文章占有效作品的比例
      
          阅读量>=100001 即微信生态的"10万+"爆款(100001为截断下界,真实值更高)。
          """
          if not works:
              return 0
          valid = [_work_read(w) for w in works if _work_read(w) > 0]
          if not valid:
              return 0
          viral_count = sum(1 for r in valid if r >= 100001)
          return viral_count / len(valid)
      
      
      # ════════════════════════════════════════════════════════════
      #  四维度评分函数
      # ════════════════════════════════════════════════════════════
      
      def _score_content_health(works, signature, verify_name, account_type):
          """内容健康度评分(原始分0-10分,综合评分时由analyzer按 原始分/10×100 转为百分制,权重按分类自适应,默认25%)
      
          更新稳定性(15%): 10分=日更,8分=周更3-5次,5分=周更1-2次,<5=不规律
          内容垂直度(15%): 10分=极度垂直单一领域,7分=主领域+偶尔跨界,<5=内容杂乱
          原创能力(10%): 10分=90%+原创,8分=70-90%原创,5分=50-70%,<5=大量转载
          质量稳定性(10%): 阅读量波动系数(排除10万+截断值后计算),波动≤35%得高分;全部10万+直接满分
          内容深度(5%): 长文比例、专业引用、独家观点
          形式创新(5%): 多媒体运用、互动形式、排版创新
          总权重=60%→归一化为0-10分→analyzer转换为百分制(综合评分权重25%)
          """
          if not works:
              return {
                  "更新稳定性": 0, "内容垂直度": 0, "原创能力": 0,
                  "质量稳定性": 0, "内容深度": 0, "形式创新": 0,
                  "原始分": 0, "总分": 0
              }
      
          # 更新稳定性(15%): 7天发文>=5篇满分, 3-4篇0.7, 1-2篇0.3, 0篇0
          work_count = len(works)
          if work_count >= 5:
              update_stability = 1.0
          elif work_count >= 3:
              update_stability = 0.7
          elif work_count >= 1:
              update_stability = 0.3
          else:
              update_stability = 0
      
          # 内容垂直度(15%): 基于accountType和works标题关键词匹配
          # accountType是平台分类标签(如"人文资讯"),存在即代表平台已确认账号领域定位,
          # 给基础分0.75;标题命中分类关键词时作为增强,避免整串匹配失效导致误判
          vertical_score = 0.75  # 有分类标签的基础分(公众号通常有明确定位)
          if account_type and works:
              type_keywords = set(account_type.split())
              match_count = 0
              for w in works:
                  title = (w.get("title") or "").lower()
                  if any(kw in title for kw in type_keywords):
                      match_count += 1
              match_ratio = match_count / len(works) if works else 0
              if match_ratio >= 0.9:
                  vertical_score = 1.0
              elif match_ratio >= 0.7:
                  vertical_score = 0.85
              elif match_ratio >= 0.4:
                  vertical_score = 0.75
              # match_ratio < 0.4 时维持基础分:分类标签存在即有明确定位
      
          # 原创能力(10%): 综合多种原创标记字段判断
          # 优先查works中的原创字段,其次查标题文字,无法判断时给中等分(不惩罚优质未标记内容)
          original_count = sum(
              1 for w in works
              if w.get("isOriginal") or w.get("originalFlag") or w.get("type") == "original"
              or "原创" in (w.get("title") or "")
          )
          original_ratio = original_count / len(works) if works else 0
          if original_ratio >= 0.7:
              original_score = 1.0
          elif original_ratio >= 0.4:
              original_score = 0.75
          elif original_ratio > 0:
              original_score = 0.70
          else:
              # 无原创标记:不代表非原创,给较高中等分(0.75),避免错误惩罚优质账号
              original_score = 0.75
      
          # 质量稳定性(10%): 基于阅读量变异系数(标准差/均值)
          # 排除10万+截断值(100001):该值非真实阅读量,计入会虚增波动、反向惩罚爆款账号;
          # 全部为10万+的账号直接给满分(常态爆款即顶级稳定)
          valid_reads_all = [_work_read(w) for w in works if _work_read(w) > 0]
          viral_all = bool(valid_reads_all) and all(r >= 100001 for r in valid_reads_all)
          if viral_all:
              quality_stability = 1.0
              cv = 0
          else:
              reads = [_work_read(w) for w in works if 0 < _work_read(w) < 100001]
              if len(reads) >= 3:
                  import statistics
                  mean_read = statistics.mean(reads)
                  std_read = statistics.stdev(reads)
                  cv = std_read / mean_read if mean_read > 0 else 1
                  if cv <= 0.35:
                      quality_stability = 1.0
                  elif cv <= 0.6:
                      quality_stability = 0.6
                  elif cv <= 0.9:
                      quality_stability = 0.3
                  else:
                      quality_stability = 0.1
              else:
                  quality_stability = 0.5
                  cv = 0
      
          # 内容深度(5%): 基于标题长度(长标题通常信息更丰富)
          avg_title_len = sum(len(w.get("title") or "") for w in works) / len(works) if works else 0
          if avg_title_len >= 20:
              depth_score = 1.0
          elif avg_title_len >= 14:
              depth_score = 0.7
          elif avg_title_len >= 8:
              depth_score = 0.4
          else:
              depth_score = 0.2
      
          # 形式创新(5%): 基于封面图多样性(有coverUrl的比例)
          covers = [w for w in works if w.get("coverUrl")]
          cover_ratio = len(covers) / len(works) if works else 0
          innovation_score = min(1.0, cover_ratio * 1.2)
      
          # 加权计算原始分(0-10)
          raw_score = (
              update_stability * 0.15 +
              vertical_score * 0.15 +
              original_score * 0.10 +
              quality_stability * 0.10 +
              depth_score * 0.05 +
              innovation_score * 0.05
          ) / 0.60 * 10  # 归一化到0-10
      
          raw_score = round(min(10, max(0, raw_score)), 1)
      
          return {
              "更新稳定性": round(update_stability * 10, 1),
              "内容垂直度": round(vertical_score * 10, 1),
              "原创能力": round(original_score * 10, 1),
              "质量稳定性": round(quality_stability * 10, 1),
              "内容深度": round(depth_score * 10, 1),
              "形式创新": round(innovation_score * 10, 1),
              "原始分": raw_score,
              "总分": round(raw_score * 3, 1)  # 兼容字段,analyzer实际使用原始分/10×100转百分制
          }
      
      
      def _score_user_activity(works, interaction_rate=0, category="general"):
          """用户活跃度评分(原始分0-10分,综合评分时由analyzer按 原始分/10×100 转为百分制,权重按分类自适应)
      
          v4.1: 根据账号分类自动调整互动率/评论密度/分享率的评分阈值
          互动率(20%): (点赞+在看+留言)/阅读量
          留言质量(10%): 评论密度(公众号评论需审核)
          分享传播力(10%): 分享率
          阅读完成率(5%): 基于互动率+绝对互动量推断
          活跃时段集中度(5%): 推送时间固定性
          总权重=50%→归一化为0-10分→analyzer转换为百分制
          """
          if not works:
              return {
                  "互动率": 0, "留言质量": 0, "分享传播力": 0,
                  "阅读完成率": 0, "活跃时段集中度": 0,
                  "原始分": 0, "总分": 0,
                  "_extra": {"interaction_rate": 0, "comment_density": 0, "category": category}
              }
      
          # 计算各项指标
          total_reads = sum(_work_read(w) for w in works)
          total_likes = sum(_work_like(w) for w in works)
          total_comments = sum(_work_comment(w) for w in works)
          total_shares = sum(_work_share(w) for w in works)
          total_watches = sum(w.get("watchCount") or 0 for w in works)
      
          # v4.1: 获取分类对应的阈值配置
          thresholds = CATEGORY_ACTIVITY_THRESHOLDS.get(category, CATEGORY_ACTIVITY_THRESHOLDS["general"])
          inter_thresh = thresholds["interaction"]
          comment_thresh = thresholds["comment"]
          share_thresh = thresholds["share"]
      
          # 互动率(20%): (点赞+评论+分享+在看)/阅读数 — 分类自适应阈值
          if total_reads > 0:
              inter_rate = (total_likes + total_comments + total_shares + total_watches) / total_reads
          else:
              inter_rate = 0
          if inter_rate >= inter_thresh[1.0]:
              interaction_score = 1.0
          elif inter_rate >= inter_thresh[0.7]:
              interaction_score = 0.7
          elif inter_rate >= inter_thresh[0.4]:
              interaction_score = 0.4
          else:
              interaction_score = 0.1
      
          # 留言质量(10%): 评论数/阅读数 — 分类自适应阈值
          # 公众号评论需作者审核,评论密度远低于开放平台
          comment_density = total_comments / total_reads if total_reads > 0 else 0
          if comment_density >= comment_thresh[1.0]:
              comment_score = 1.0
          elif comment_density >= comment_thresh[0.7]:
              comment_score = 0.7
          elif comment_density >= comment_thresh[0.4]:
              comment_score = 0.4
          else:
              comment_score = 0.1
      
          # 分享传播力(10%): 分享数/阅读数 — 分类自适应阈值
          share_rate = total_shares / total_reads if total_reads > 0 else 0
          if share_rate >= share_thresh[1.0]:
              share_score = 1.0
          elif share_rate >= share_thresh[0.7]:
              share_score = 0.7
          elif share_rate >= share_thresh[0.4]:
              share_score = 0.4
          else:
              share_score = 0.1
      
          # 阅读完成率(5%): 无法精确获取,根据互动率+绝对互动量推断
          # 大号绝对互动量极高(100万+总互动),阅读完成率不会低于0.6
          read_completion = min(1.0, inter_rate * 8) if inter_rate > 0 else 0.3
          if total_reads > 0:
              total_interactions = total_likes + total_comments + total_shares + total_watches
              if total_interactions >= 100000:
                  read_completion = max(read_completion, 0.6)
              if total_interactions >= 500000:
                  read_completion = max(read_completion, 0.7)
      
          # 活跃时段集中度(5%): 发布时间规律性
          from collections import Counter
          hours = []
          for w in works:
              pub_time = w.get("publishTime", "") or ""
              if pub_time:
                  try:
                      if isinstance(pub_time, (int, float)):
                          if pub_time > 1e12:
                              pub_time = pub_time / 1000
                          hour = datetime.fromtimestamp(pub_time).hour
                      else:
                          hour = int(str(pub_time)[11:13]) if len(str(pub_time)) > 13 else -1
                      if hour >= 0:
                          hours.append(hour)
                  except (ValueError, OSError):
                      pass
      
          if hours:
              hour_counter = Counter(hours)
              top_hour_count = hour_counter.most_common(1)[0][1]
              concentration = top_hour_count / len(hours)
              if concentration >= 0.6:
                  time_score = 1.0
              elif concentration >= 0.4:
                  time_score = 0.6
              else:
                  time_score = 0.3
          else:
              time_score = 0.3
      
          # 加权计算原始分(0-10)
          raw_score = (
              interaction_score * 0.20 +
              comment_score * 0.10 +
              share_score * 0.10 +
              read_completion * 0.05 +
              time_score * 0.05
          ) / 0.50 * 10  # 归一化到0-10
      
          raw_score = round(min(10, max(0, raw_score)), 1)
      
          return {
              "互动率": round(interaction_score * 10, 1),
              "留言质量": round(comment_score * 10, 1),
              "分享传播力": round(share_score * 10, 1),
              "阅读完成率": round(read_completion * 10, 1),
              "活跃时段集中度": round(time_score * 10, 1),
              "原始分": raw_score,
              "总分": round(raw_score * 2.5, 1),  # 兼容字段,analyzer实际使用原始分/10×100转百分制
              "_extra": {
                  "interaction_rate": round(inter_rate * 100, 2),
                  "comment_density": round(comment_density * 100, 4),
                  "category": category,
              }
          }
      
      
      def _score_core_data(works, avg_read_count, category="general"):
          """内容核心数据表现评分(0-31分制,综合评分时转换为百分制)
      
          v4.1: 根据账号分类调整互动类子项的评分灵敏度
          阅读数表现(8分): 平均阅读数区间分段
          点赞数表现(6分): 平均点赞数+点赞率
          评论数表现(4分): 平均评论数+评论率
          互动率表现(3分): 综合互动率
          发布时间合理性(2分): 黄金时段发文比例
          爆款产出力(8分): 10万+阅读文章占比(爆款率)
          满分=31分,转换为百分制:raw/31*100
          """
          if not works:
              return {
                  "阅读数表现": 0, "点赞数表现": 0, "评论数表现": 0,
                  "互动率表现": 0, "发布时间合理性": 0,
                  "爆款产出力": 0,
                  "原始分": 0, "总分": 0
              }
      
          # 1. 阅读数表现(8分) - 基于公众号行业真实水平校准
          # 行业参考:普通账号500-3000,良好账号3000-1万,优质账号1-5万,顶尖账号5万+
          avg_read = avg_read_count or 0
          if avg_read >= 100000:
              read_score = 8    # 超顶级,如人民日报等媒体大号
          elif avg_read >= 80000:
              read_score = 7.5  # 行业顶尖(8万+)
          elif avg_read >= 50000:
              read_score = 6.5  # 行业顶尖(5万+)
          elif avg_read >= 30000:
              read_score = 5    # 行业优秀(3-5万)
          elif avg_read >= 20000:
              read_score = 4    # 行业优秀(2-3万)
          elif avg_read >= 10000:
              read_score = 3    # 行业良好(1-2万)
          elif avg_read >= 5000:
              read_score = 2    # 中等水平(5000-1万)
          elif avg_read >= 1000:
              read_score = 1    # 偏低
          else:
              read_score = 0.5  # 极低
      
          # 2. 点赞数表现(6分)
          total_likes = sum(_work_like(w) for w in works)
          total_reads = sum(_work_read(w) for w in works)
          avg_likes = total_likes / len(works) if works else 0
          like_rate = total_likes / total_reads if total_reads > 0 else 0
      
          if avg_likes >= 3000 or like_rate >= 0.03:
              like_score = 6
          elif avg_likes >= 1000 or like_rate >= 0.015:
              like_score = 5
          elif avg_likes >= 300 or like_rate >= 0.005:
              like_score = 3
          elif avg_likes >= 100 or like_rate >= 0.003:
              like_score = 1.5
          else:
              like_score = 1
      
          # 3. 评论数表现(4分)
          # 微信公众号评论生态校准:评论需经作者审核展示,大号篇均5-30条已属优秀,
          # 旧阈值(500+/200+/100+)对标开放平台(抖音/小红书),严重低估公众号头部账号
          total_comments = sum(_work_comment(w) for w in works)
          avg_comments = total_comments / len(works) if works else 0
          comment_rate = total_comments / total_reads if total_reads > 0 else 0
      
          if avg_comments >= 20 or comment_rate >= 0.0005:
              comment_score = 4
          elif avg_comments >= 10 or comment_rate >= 0.0003:
              comment_score = 3
          elif avg_comments >= 5 or comment_rate >= 0.0001:
              comment_score = 2
          elif avg_comments >= 2:
              comment_score = 1
          else:
              comment_score = 0.5
      
          # 4. 互动率表现(3分) - 基于公众号行业真实水平校准
          # 行业参考:普通1-3%,良好3-8%,优秀8-15%,顶尖15%+
          total_shares = sum(_work_share(w) for w in works)
          total_watches = sum(w.get("watchCount") or 0 for w in works)
          inter_rate = (total_likes + total_comments + total_shares + total_watches) / total_reads if total_reads > 0 else 0
      
          if inter_rate >= 0.10:    # 顶尖:10%+(极少账号能达到)
              inter_score = 3
          elif inter_rate >= 0.05:   # 优秀:5-10%
              inter_score = 2.5
          elif inter_rate >= 0.03:  # 良好:3-5%
              inter_score = 2
          elif inter_rate >= 0.015: # 中等:1.5-3%
              inter_score = 1.5
          elif inter_rate >= 0.005:  # 偏低:0.5-1.5%
              inter_score = 1
          else:
              inter_score = 0
      
          # 5. 发布时间合理性(2分): 黄金时段发文比例
          golden_count = 0
          total_with_time = 0
          for w in works:
              pub_time = w.get("publishTime", "") or ""
              hour = -1
              if pub_time:
                  try:
                      if isinstance(pub_time, (int, float)):
                          if pub_time > 1e12:
                              pub_time = pub_time / 1000
                          hour = datetime.fromtimestamp(pub_time).hour
                      else:
                          hour = int(str(pub_time)[11:13]) if len(str(pub_time)) > 13 else -1
                  except (ValueError, OSError):
                      pass
              if hour >= 0:
                  total_with_time += 1
                  if (7 <= hour <= 9) or (12 <= hour <= 13) or (20 <= hour <= 22):
                      golden_count += 1
      
          if total_with_time > 0:
              golden_ratio = golden_count / total_with_time
              if golden_ratio >= 0.6:
                  time_score = 2
              elif golden_ratio >= 0.3:
                  time_score = 1.5
              elif golden_ratio >= 0.1:
                  time_score = 1
              else:
                  time_score = 0
          else:
              time_score = 0.5
      
          # 6. 爆款产出力(8分): 10万+阅读文章占比(微信生态顶级传播力信号)
          # 10万+是微信阅读量显示上限,爆款率>20%即为头部账号,40%+为超级头部
          viral_ratio = _calc_viral_ratio(works)
          if viral_ratio >= 0.40:
              viral_score = 8
          elif viral_ratio >= 0.30:
              viral_score = 7
          elif viral_ratio >= 0.25:
              viral_score = 6.5
          elif viral_ratio >= 0.20:
              viral_score = 6
          elif viral_ratio >= 0.10:
              viral_score = 4.5
          elif viral_ratio >= 0.05:
              viral_score = 3
          elif viral_ratio >= 0.02:
              viral_score = 2
          elif viral_ratio > 0:
              viral_score = 1
          else:
              viral_score = 0
      
          # v4.1: 分类感知的互动增益
          # 时政/新闻号互动天然低(读者被动消费),对互动类子项给予分类增益
          # 增益后不超过各子项满分上限
          _engagement_boost = {
              "news_politics": 2.0,     # 时政号互动分数×2(上限封顶)
              "finance_opinion": 1.5,   # 财经号互动分数×1.5
          }
          boost = _engagement_boost.get(category, 1.0)
          if boost > 1.0:
              like_score = min(6, like_score * boost)
              comment_score = min(4, comment_score * boost)
              inter_score = min(3, inter_score * boost)
      
          raw_score = round(
              read_score + like_score + comment_score + inter_score + time_score
              + viral_score, 1
          )
      
          return {
              "阅读数表现": read_score,
              "点赞数表现": like_score,
              "评论数表现": comment_score,
              "互动率表现": inter_score,
              "发布时间合理性": time_score,
              "爆款产出力": viral_score,
              "爆款率": round(viral_ratio * 100, 1),
              "原始分": raw_score,
              "总分": round(raw_score / 31 * 100, 1)  # 转换为百分制
          }
      
      
      def _score_operation_compliance(works, verify_name):
          """运营规范性评分(直接0-10分,综合评分权重按分类自适应,默认15%)
      
          更新频率(5分): 7天发文数
          发布时间合理性(3分): 固定时段发文比例
          账号认证(2分): 是否有认证
          """
          if not works:
              return {
                  "更新频率": 0, "发布时间合理性": 0, "账号认证": 0,
                  "原始分": 0
              }
      
          # 更新频率(5分)
          work_count = len(works)
          if work_count >= 5:
              freq_score = 5
          elif work_count >= 3:
              freq_score = 4
          elif work_count >= 2:
              freq_score = 3
          elif work_count >= 1:
              freq_score = 2
          else:
              freq_score = 0
      
          # 发布时间合理性(3分): 固定时段发文占比
          hours = []
          for w in works:
              pub_time = w.get("publishTime", "") or ""
              if pub_time:
                  try:
                      if isinstance(pub_time, (int, float)):
                          if pub_time > 1e12:
                              pub_time = pub_time / 1000
                          hour = datetime.fromtimestamp(pub_time).hour
                      else:
                          hour = int(str(pub_time)[11:13]) if len(str(pub_time)) > 13 else -1
                      if hour >= 0:
                          hours.append(hour)
                  except (ValueError, OSError):
                      pass
      
          if hours:
              from collections import Counter
              hour_counter = Counter(hours)
              top_hour_count = hour_counter.most_common(1)[0][1]
              regularity = top_hour_count / len(hours)
              if regularity >= 0.6:
                  time_score = 3
              elif regularity >= 0.4:
                  time_score = 2
              elif regularity >= 0.2:
                  time_score = 1
              else:
                  time_score = 0
          else:
              time_score = 0.5
      
          # 账号认证(2分)
          auth_score = 2 if verify_name else 0
      
          raw_score = round(freq_score + time_score + auth_score, 1)
      
          return {
              "更新频率": freq_score,
              "发布时间合理性": time_score,
              "账号认证": auth_score,
              "原始分": raw_score
          }
      
      
      # ════════════════════════════════════════════════════════════
      #  等级判定
      # ════════════════════════════════════════════════════════════
      
      def _get_score_level(score, max_score):
          """根据得分率返回评级(含图标):优/良/中/差"""
          rate = score / max_score * 100 if max_score > 0 else 0
          if rate >= 80:
              return "优"
          elif rate >= 60:
              return "良"
          elif rate >= 40:
              return "中"
          else:
              return "差"
      
      
      def _get_score_level_icon(score, max_score):
          """根据得分率返回评级+图标"""
          rate = score / max_score * 100 if max_score > 0 else 0
          if rate >= 80:
              return "优 ⭐"
          elif rate >= 60:
              return "良 ✅"
          elif rate >= 40:
              return "中 📊"
          else:
              return "差 ⚠️"
      
      
      def _get_overall_grade(score):
          """根据综合评分返回等级图标+评级+等级
      
          评级标准(基于公众号行业真实水平校准):
          公众号生态中均阅3万+即属top 0.1%头部,80分+即为S级标杆
          S级(行业标杆) >=80:均阅3万+、稳定日更、高互动的顶尖账号
          A级(优质账号) >=70:均阅1-3万、持续优质内容的优质账号
          B级(健康账号) >=60:正常运营、有稳定输出的健康账号
          C级(中等账号) >=50:基础运营待提升
          D级(亚健康)   >=40:运营不稳定或互动偏低
          E级(问题账号)  <40:严重问题需全面诊断
          """
          if score >= 80:
              return "🏆 标杆账号", "S级"
          elif score >= 70:
              return "⭐ 优质账号", "A级"
          elif score >= 60:
              return "✅ 健康账号", "B级"
          elif score >= 50:
              return "📊 中等账号", "C级"
          elif score >= 40:
              return "⚠️ 亚健康账号", "D级"
          else:
              return "❌ 问题账号", "E级"
      
    • wechat_analyzer.py 3.7 KB
      """wechat_analyzer.py — 公众号账号诊断入口点
      
      模块化拆分后,本文件仅负责 argparse 命令路由。
      核心逻辑分布在以下模块中:
        - api_client.py  — 红狐 API HTTP 通信与凭证管理
        - scoring.py     — 作品辅助函数、四维度评分、等级判定
        - analyzer.py    — 单账号分析编排、查询/同步命令
        - report.py      — HTML 报告模板替换与生成
      
      向后兼容:scoring 模块的公开函数通过 __all__ 重导出,
      使 `from wechat_analyzer import _work_read` 等旧导入继续可用。
      """
      
      import argparse
      import json
      import sys
      
      from analyzer import cmd_query, cmd_sync_notes
      from report import cmd_generate_html, cmd_generate_multi_html
      
      # ── 向后兼容重导出(供 test_wechat_analyzer.py 等旧导入使用)──
      from scoring import (  # noqa: F401
          _work_read,
          _work_like,
          _work_comment,
          _work_share,
          _work_collect,
          _work_interact_total,
          _work_publish_time,
          _calc_avg_read,
          _calc_viral_ratio,
          _score_content_health,
          _score_user_activity,
          _score_core_data,
          _score_operation_compliance,
      )
      
      
      def main():
          parser = argparse.ArgumentParser(description="公众号账号诊断宗师")
          subparsers = parser.add_subparsers(dest="command")
      
          # 查询子命令
          query_parser = subparsers.add_parser("query", help="查询账号数据")
          query_parser.add_argument("--account_ids", help="公众号账号ID列表,逗号分隔", required=False)
          query_parser.add_argument("--account_names", help="公众号账号名称列表,逗号分隔", required=False)
          query_parser.add_argument("--force_analyze", action="store_true", help="强制执行分析(即使无作品数据)")
      
          # 同步作品子命令
          sync_parser = subparsers.add_parser("sync_notes", help="同步账号作品数据")
          sync_parser.add_argument("--account_id", help="公众号账号ID(单个)", required=True)
          sync_parser.add_argument("--account_names", help="账号名称列表,逗号分隔(用于提示)", required=False)
      
          # 生成单账号HTML子命令
          html_parser = subparsers.add_parser("generate_html", help="基于report_data.json生成单账号HTML报告")
      
          # 生成多账号对比HTML子命令
          multi_parser = subparsers.add_parser("generate_multi_html", help="基于multi_report_data.json生成多账号对比HTML报告")
      
          args = parser.parse_args()
      
          if args.command == "query":
              account_ids = [x.strip() for x in args.account_ids.split(",") if x.strip()] if args.account_ids else []
              account_names = [x.strip() for x in args.account_names.split(",") if x.strip()] if args.account_names else []
      
              if not account_ids and not account_names:
                  print(json.dumps({
                      "status": "error",
                      "message": "请提供至少一个账号ID(--account_ids)或名称(--account_names)"
                  }, ensure_ascii=False))
                  sys.exit(1)
      
              cmd_query(account_ids=account_ids or None, account_names=account_names or None, force_analyze=getattr(args, 'force_analyze', False))
      
          elif args.command == "sync_notes":
              account_id = args.account_id.strip() if args.account_id else ""
              if not account_id:
                  print(json.dumps({
                      "status": "error",
                      "message": "请提供账号ID(--account_id)"
                  }, ensure_ascii=False))
                  sys.exit(1)
              cmd_sync_notes([account_id])
      
          elif args.command == "generate_html":
              cmd_generate_html()
      
          elif args.command == "generate_multi_html":
              cmd_generate_multi_html()
      
          else:
              parser.print_help()
              sys.exit(1)
      
      
      if __name__ == "__main__":
          main()
      
  • README.en.md 3.3 KB
    # WeChat Official Account Analyzer / wechat-account-analyzer
    
    ---
    
    ## Overview
    
    WeChat Official Account Analyzer is an intelligent diagnostic tool. Simply enter an account name to receive a four-dimensional quantitative scoring report, benchmark against industry averages, and get actionable optimization suggestions.
    
    **Core Value**
    
    - 🎯 **Data-Driven Decisions**: Automatic scoring based on real operational data — no more guesswork; see the full account picture in one diagnosis
    - 📊 **Four-Dimensional Diagnosis**: Content health, user engagement, core metrics, and operational compliance — covering all key dimensions
    - 💡 **Actionable Suggestions**: Prioritized as urgent / key / continuous, each with expected quantitative results to guide your next move
    - 🔍 **Industry Benchmarking**: Auto-match peer accounts for horizontal comparison and clarify your competitive position
    
    **Target Users**
    
    - 👤 Account Owners — Understand account health and find optimization directions
    - 📱 Social Media Operators — Data-driven strategy with measurable optimization results
    - 🏢 MCN Agencies — Evaluate account value to support partnership decisions
    - 🏷️ Brands & Content Creators — Competitive analysis and best practice learning
    
    ---
    
    ## Features
    
    ### Core Capabilities
    
    - 📊 **Four-Dimensional Scoring**: Diagnose content health / user engagement / core data / operational compliance, with v4.1 category-adaptive weighting (auto-tuned weights/thresholds for 6 account types) and explicit viral-hit (100k+ reads) scoring
    - 🏆 **Smart Rating**: S/A/B/C/D/E six-level rating (S ≥ 80) with industry benchmarking to understand your standing
    - 📈 **Content Data Display**: Views, likes, comments, and reposts for recent posts at a glance
    - 💡 **Optimization Suggestions**: Prioritized with "Problem → Suggestion → Expected Result" structure
    - 🔍 **Similar Account Discovery**: Auto-match peer accounts for horizontal comparison
    - 🔗 **One-Click Navigation**: Account names and post titles link directly to WeChat for details
    
    ---
    
    ## Usage Guide
    
    Describe your needs in natural language — no commands to memorize.
    
    ### Quick Reference
    
    | Intent | Example | Result |
    |--------|---------|--------|
    | Diagnose an account | "Diagnose account-name" | Full diagnostic report |
    | Compare accounts | "Compare A and B" | Side-by-side comparison + insights |
    | Query by ID | "Diagnose account ID xxx" | Precise lookup by ID |
    
    ### Output Example
    
    Reports follow a fixed 5-chapter structure:
    
    > **1. Account Info** → Name, ID, verified entity, description
    > **2. Overall Score** → 100-point score + industry benchmarking table
    > **3. Recent Posts** → Detailed metrics for recent posts
    > **4. Optimization Suggestions** → Urgent / Key / Continuous actionable advice
    > **5. Industry Benchmarking** → Benchmark conclusions + similar account recommendations
    
    ---
    
    ## Use Cases
    
    | Scenario | Role | Example Query | Benefit |
    |----------|------|---------------|---------|
    | Self-Audit | Account Owner | "Diagnose my account" | Identify health status and quantify improvements |
    | Competitor Analysis | Operator | "Analyze account X" | Track competitors and learn from winners |
    | Portfolio Evaluation | MCN Agency | "Evaluate account X" | Data-driven partnership decisions |
    | Content Strategy | Content Creator | "Diagnose and suggest direction" | Evidence-based content planning |
    
  • README.md 3 KB
    # 公众号账号诊断 / wechat-account-analyzer
    
    ---
    
    ## 简介
    
    公众号账号诊断是一款智能分析工具,输入公众号名称即可获取四维度量化评分报告,对标行业平均水平,输出可落地的运营优化建议。
    
    **核心价值**
    
    - 🎯 **数据驱动决策**:基于真实运营数据自动评分,告别主观判断,一次诊断即可看清账号全貌
    - 📊 **四维全面诊断**:内容健康度、用户活跃度、核心数据、运营规范性,覆盖运营关键维度
    - 💡 **可执行建议**:分紧急/重点/持续三级,每条建议含预期量化效果,直接指导优化动作
    - 🔍 **对标行业水平**:自动匹配同赛道账号,横向对比找差距,明确自身行业位置
    
    **适用对象**
    
    - 👤 公众号号主 — 了解账号健康度,找到优化方向
    - 📱 新媒体运营 — 数据驱动运营策略,量化优化效果
    - 🏢 MCN 机构 — 评估账号价值,辅助签约决策
    - 🏷️ 品牌方与内容创作者 — 竞品分析,借鉴成功经验
    
    ---
    
    ## 功能特性
    
    ### 核心功能
    
    - 📊 **四维度评分**:内容健康度 / 用户活跃度 / 核心数据 / 运营规范性四维度诊断,v4.1 分类自适应(6 类账号自动调整权重/阈值),爆款率(10万+)显式计分
    - 🏆 **智能评级**:S/A/B/C/D/E 六级评级(S≥80 分)+ 行业对标分析,直观了解账号段位
    - 📈 **作品数据展示**:近期作品阅读量、点赞数、评论数、在看数一目了然
    - 💡 **优化建议**:按紧急程度分级,每条采用"问题 → 建议 → 预期"三段结构
    - 🔍 **相似账号发现**:自动匹配同赛道对标账号,横向对比找差距
    - 🔗 **一键跳转**:账号名称、作品标题均支持跳转微信查看详情
    
    ---
    
    ## 使用指南
    
    直接用自然语言描述需求,无需记忆命令。
    
    ### 常用说法速查
    
    | 意图 | 示例话术 | 效果 |
    |------|---------|------|
    | 诊断公众号 | "诊断十点读书" | 输出完整诊断报告 |
    | 多账号对比 | "对比诊断 A 和 B" | 横向对比 + 差异化建议 |
    | 按 ID 查询 | "诊断公众号ID xxx" | 通过 ID 精确查询 |
    
    ### 输出示例
    
    诊断报告以固定的 5 章节结构输出:
    
    > **一、账号信息** → 名称、ID、认证主体、账号简介
    > **二、综合评分** → 百分制评分 + 行业对标数据表
    > **三、近期作品数据** → 最近作品详细数据表格
    > **四、优化建议** → 紧急 / 重点 / 持续三级可执行建议
    > **五、行业对标分析** → 对标结论 + 相似账号推荐
    
    ---
    
    ## 使用场景
    
    | 场景 | 角色 | 示例问法 | 收益 |
    |------|------|---------|------|
    | 号主自检优化 | 公众号号主 | "诊断我的公众号" | 明确账号健康度,量化优化效果 |
    | 竞品对标分析 | 运营人员 | "分析 XX 公众号" | 掌握竞品动态,借鉴成功经验 |
    | MCN 批量评估 | MCN 机构 | "评估 XX 账号价值" | 数据驱动签约决策 |
    | 内容策略制定 | 内容创作者 | "诊断并分析内容方向" | 有据可依的内容规划 |
    
  • SKILL.md 10 KB
    ---
    name: wechat-account-analyzer
    description: 公众号账号诊断工具是对任意公众号账号进行四维度量化评分(内容健康度、用户活跃度、内容核心数据、运营规范性),对标行业平均水平,输出可落地的运营优化建议。
    ---
    
    # 公众号账号诊断
    
    > 基于多维度数据分析,为公众号提供全面的账号健康诊断和优化建议
    
    ---
    
    ## 简介
    
    **公众号账号诊断** 是一款智能分析工具,通过红狐 API 获取公众号真实运营数据,基于内容健康度、用户活跃度、核心数据表现、运营规范性四维度进行评分诊断,自动生成对标行业水平的优化建议。
    
    - **核心价值**:告别主观判断,用数据驱动公众号运营决策。一次诊断即可看清账号全貌,明确优化方向
    - **适用对象**:公众号号主、新媒体运营、MCN机构、品牌方、内容创作者
    - **技术基础**:基于红狐API实时数据 + Python评分引擎 + Agent智能建议生成
    
    ---
    
    ## 功能特性
    
    ### 核心能力
    
    | 功能 | 说明 |
    |------|------|
    | 📊 **四维度评分** | 四维度诊断评分(内容健康度/用户活跃度/核心数据/运营规范性),v4.1分类自适应(6类账号自动调整权重/阈值),爆款率(10万+)显式计分 |
    | 🏆 **智能评级** | S/A/B/C/D/E 六级评级(S≥80分)+ 行业对标分析 |
    | 📈 **数据可视化** | 近期作品数据表格,阅读/点赞/评论/在看一目了然 |
    | 💡 **优化建议** | 分紧急/重点/持续三级,每条建议含预期量化效果 |
    | 🔍 **相似账号** | 自动匹配同赛道对标账号,横向对比找差距 |
    | 🔗 **可跳转链接** | 账号名称、作品标题均支持一键跳转微信 |
    
    ### 特色亮点
    
    - **⚡ 一句话诊断**:输入公众号名称即可获取完整报告,无需复杂配置
    - **🔄 多账号对比**:支持同时诊断多个账号并生成对比报告
    - **📋 强制格式输出**:5章节固定模板,章节/顺序/格式强制锁定,不得偏离
    
    ---
    
    ## 一键安装
    
    ### 前置条件
    - 已安装 Python 3.8+
    - 已注册 [红狐Hub](https://redfox.hk/) 账号并获取 API Key
    
    ### 安装步骤
    
    1. 将技能文件夹放入你的 Skills 目录
    2. 安装 Python 依赖:
       ```bash
       pip install requests
       ```
    3. 配置 API Key(见下方)
    
    ### 配置 API Key
    
    #### 获取 API Key
    1. 访问 [红狐Hub 官网](https://redfox.hk/) 了解服务详情
    2. 前往 [注册页面](https://redfox.hk/login) 注册账号
    3. **新注册用户将获赠免费积分**,可立即开始使用
    4. 注册登录后,在个人中心获取 API Key,格式为 `ak_xxxxxxxx`
    
    #### 设置环境变量
    
    `REDFOX_API_KEY` 需从环境变量获取。若未设置,Agent 会主动帮你配置:
    
    - **macOS/Linux**:将 `export REDFOX_API_KEY=<值>` 追加到 `~/.zshrc` 或 `~/.bashrc`,然后 `source` 对应文件
    - **Windows**:使用 `[Environment]::SetEnvironmentVariable("REDFOX_API_KEY", "<值>", "User")` 设置用户级永久环境变量(需重启终端)
    
    配置完成后验证:`echo $REDFOX_API_KEY`(macOS/Linux)或 `echo %REDFOX_API_KEY%`(Windows)
    
    ---
    
    ## 使用指南
    
    ### 触发方式
    
    当用户提到以下任意关键词时自动激活:
    - "诊断公众号"、"账号分析"、"公众号体检"
    - "账号评估"、"查看XX公众号数据"、"分析XX账号"
    - 直接输入公众号名称 + "诊断/分析"(如"十点读书 诊断")
    
    首次交互时 Agent 会主动打招呼并引导用户输入账号。
    
    ### 常用命令速查
    
    | 命令 | 说明 |
    |------|------|
    | `诊断"公众号名称"` | 输入名称即可获取完整诊断报告 |
    | `对比诊断"A"和"B"` | 多账号横向对比 + 差异化建议 |
    | `诊断公众号ID"xxx"` | 通过公众号ID查询(非中文名) |
    
    ### 使用示例
    
    **示例1:标准诊断**
    > 用户:诊断"十点读书"公众号
    > AI:输出完整诊断报告(账号信息→综合评分→作品数据→优化建议→行业对标)
    
    **示例2:多账号对比**
    > 用户:对比诊断"十点读书"和"洞见"
    > AI:分别输出两个诊断报告 + 横向对比总结表格 + 差异化建议
    
    **示例3:未查到账号**
    > 用户:诊断“不存在的公众号123”
    > AI:未查询到账号【不存在的公众号123】,该账号可能尚未收录或名称有误,请核实公众号名称后重试。
    
    ### ⚠️ 输出格式强制规范
    
    诊断报告**必须且仅能**以5章节固定模板输出:**一、账号信息 → 二、综合评分 → 三、近期作品数据 → 四、优化建议 → 五、行业对标分析**。章节不可省略、顺序不可调换、格式不可偏离。所有字段名、表格结构、emoji 标识须与模板严格一致。违反模板的输出视为不合格。
    
    > 详细技术规范、评分算法、输出格式模板请参见 **[核心工作流](references/core_workflow.md)**
    
    ---
    
    ## 使用场景
    
    ### 场景一:号主自检优化
    **需求**:了解自己公众号在行业中的水平,找到优化方向
    
    **使用方式**:
    1. 直接输入自己的公众号名称进行诊断
    2. 查看四维度评分找到短板
    3. 按优化建议的优先级逐步改进
    
    **预期收益**:明确账号健康度,量化优化效果,避免盲目运营
    
    ---
    
    ### 场景二:竞品对标分析
    **需求**:分析竞品账号的运营策略和数据表现
    
    **使用方式**:
    1. 诊断目标竞品公众号
    2. 查看行业对标和相似账号
    3. 对比分析差异化优势与差距
    
    **预期收益**:掌握竞品动态,借鉴成功经验,调整自身策略
    
    ---
    
    ### 场景三:MCN 批量评估
    **需求**:评估旗下或潜在签约账号的价值
    
    **使用方式**:
    1. 逐个诊断候选账号
    2. 横向对比综合评分和核心指标
    3. 结合优化建议评估成长潜力
    
    **预期收益**:数据驱动签约决策,量化账号商业价值
    
    ---
    
    ### 场景四:内容策略制定
    **需求**:根据账号诊断结果制定内容规划
    
    **使用方式**:
    1. 诊断账号获取质量稳定性、内容垂直度等评分
    2. 根据"待优化模块"确定改进重点
    3. 参考行业基准调整内容方向和发布策略
    
    **预期收益**:有据可依的内容规划,提升整体运营效率
    
    ---
    
    ## 项目架构
    
    ### 目录结构
    
    ```
    公众号账号诊断/
    ├── SKILL.md                     # 技能说明文档(本文件)
    ├── scripts/
    │   ├── wechat_analyzer.py       # 入口点:argparse 命令路由(query/sync/generate_html)
    │   ├── api_client.py             # HTTP 通信层:红狐 API 凭证管理 + POST 请求封装
    │   ├── scoring.py                # 评分引擎:作品辅助函数 + 四维度评分 + 等级判定
    │   ├── analyzer.py               # 分析编排:单账号评分处理 + 查询/同步命令
    │   ├── report.py                 # HTML 报告:模板替换 + 单账号/多账号报告生成
    │   └── test_wechat_analyzer.py   # 单元测试:辅助函数 + 评分函数容错(42 个测试)
    ├── references/
    │   ├── core_workflow.md         # 核心工作流(Agent执行参考)
    │   ├── workflow_guide.md        # 工作流详细指南
    │   └── api_guide.md             # API接口与评分逻辑说明
    ├── assets/
    │   └── report_template.html     # 单账号HTML报告模板
    └── output/                      # 输出目录(运行时自动生成)
        ├── raw_data.json            # API原始数据
        └── report_data.json         # 结构化诊断数据
    ```
    
    ### 技术栈
    
    | 组件 | 技术 |
    |------|------|
    | 运行环境 | Python 3.8+ |
    | HTTP 客户端 | requests(原生) |
    | 数据源 | 红狐API (`redfox.hk`) |
    | 输出格式 | Markdown + JSON + HTML |
    | 认证方式 | Header `X-API-KEY` + 环境变量 `REDFOX_API_KEY` |
    
    ### 核心模块
    
    | 模块文件 | 函数 | 职责 |
    |----------|------|------|
    | `wechat_analyzer.py` | `main` | argparse 入口,路由到各子命令 |
    | `api_client.py` | `https_post` | 红狐 API HTTP POST 请求封装 |
    | `api_client.py` | `_get_credential` | 环境变量 → shell 配置文件读取 API Key |
    | `scoring.py` | `_score_content_health` | 内容健康度(6子项加权,0-10分) |
    | `scoring.py` | `_score_user_activity` | 用户活跃度(5子项加权,0-10分) |
    | `scoring.py` | `_score_core_data` | 核心数据表现(6子项,0-31分) |
    | `scoring.py` | `_score_operation_compliance` | 运营规范性(3子项,0-10分) |
    | `analyzer.py` | `cmd_query` | API 查询 + 原始数据保存 |
    | `analyzer.py` | `_analyze_single_account` | 四维度评分计算 + 结构化输出 |
    | `analyzer.py` | `cmd_sync_notes` | 同步账号作品数据(订阅推送) |
    | `report.py` | `cmd_generate_html` | 单账号 HTML 报告生成 |
    | `report.py` | `cmd_generate_multi_html` | 多账号对比 HTML 报告生成 |
    
    ---
    
    ## 常见问答
    
    ### 安装相关
    
    **Q1: 提示"未找到 REDFOX_API_KEY"怎么办?**
    
    A: 请按「一键安装 → 配置 API Key」中的步骤设置环境变量。如果不会配置,直接告诉 Agent,它会帮你操作。
    
    **Q2: 需要安装什么依赖?**
    
    A: 只需 `requests` 库:`pip install requests`。Python 3.8+ 即可运行。
    
    ---
    
    ### 使用相关
    
    **Q3: 支持按公众号ID查询吗?**
    
    A: 支持。如果知道公众号ID(如 `gh_xxxxxxxx`),Agent 会自动识别并使用 ID 查询。
    
    **Q4: 可以同时诊断多个账号吗?**
    
    A: 可以。输入"对比诊断 A 和 B"即可获得多账号对比报告。
    
    **Q5: 数据多久更新一次?**
    
    A: 数据来自红狐API,时效性取决于API缓存策略。报告末尾会标注数据置信度。
    
    ---
    
    ### 故障排除
    
    **Q6: 提示"未查询到该公众号信息"?**
    
    A: 可能原因:① 账号名称输入有误;② 该账号近期发文未达收录标准。请检查名称准确性后重试。
    
    **Q7: 数据看起来不对怎么办?**
    
    A: 所有数据来自红狐API接口,诊断结果仅反映接口返回的数据。如有疑问可联系红狐平台确认数据准确性。
    
    ---
    
    ### 获取帮助
    
    - 红狐Hub 官网:[https://redfox.hk/](https://redfox.hk/)
    - 技能实现细节:参见 [核心工作流](references/core_workflow.md)
    

Comments (0)

Sign in to join the conversation.

No comments yet.

Reviews (0)

No reviews yet.

Related