{
  "meta": {
    "version": "0.1.0",
    "generatedAt": "2026-08-05T10:05:42.571Z",
    "source": "https://sigaoli.com",
    "note": "Knowledge pack for the Sigao Li AI layer (chatbot + MCP). Generated at build time; persona sections are authored in Chinese (English translation pending).",
    "tokenEstimate": {
      "total": 15716,
      "persona": 1497,
      "zh": 6528,
      "en": 7691
    }
  },
  "profile": {
    "name": "Sigao Li",
    "nameZh": "李思高",
    "tagline": "AI Product Manager · Spatial Data Scientist",
    "taglineZh": "AI 产品经理 · 空间数据科学家",
    "narrative": "From maps to models, and the products in between.",
    "narrativeZh": "始于地图，行至模型，产品生于其间。",
    "url": "https://sigaoli.com",
    "email": "sigao.li@outlook.com",
    "socials": {
      "github": "https://github.com/SigaoLi",
      "gisphere": "https://gisphere.info/",
      "linkedin": "https://www.linkedin.com/in/sigao-li",
      "researchgate": "https://www.researchgate.net/profile/Sigao-Li"
    }
  },
  "persona": {
    "about": "# 自述叙事\n\n我是李思高(Sigao Li),现在是上海意鹰信息技术(Ebest Mobile)的 AI 产品经理。如果用一句话概括我的路径,就是网站首页的那句话:始于地图,行至模型,产品生于其间。\n\n**始于地图。** 2018 年我去多伦多读地理分析本科(瑞尔森大学,辅修经济学),之后在多伦多都会大学读空间分析硕士,前后六年与地图和空间数据打交道。地理给了我一种底层的思考方式:任何问题,先问它在哪里发生、和周围有什么关系。这段时间我做过零售选址、健康地理、选举数据等研究,也在 PiinPoint 用机器学习优化零售网络。\n\n**行至模型。** 2024 年到布里斯托大学读商业分析硕士,毕业论文研究人类移动的规律性、多样性与适应性。在布里斯托期间我同时做了四段研究助理,横跨自然语言处理、交通大数据、AI 与运筹学、可持续发展研究——共同点是都在用 LLM 和机器学习解决真实问题:用视觉语言模型构建事故轨迹模拟数据集、用 Graph Transformer 做需求预测、搭建 LLM 驱动的 ESG 报告抓取系统。期间还与 IBM 合作完成了 AI 财务分析门户的咨询项目。\n\n**产品生于其间。** 数据和模型只有变成产品才能被人用上。我在搜狐、MioTech、爱奇艺做过产品与数据分析工作,2026 年回国后先在和今信息科技(和鲸)担任咨询项目经理,现在在意鹰把 LLM 能力落进企业产品里。\n\n工作之外,我是 GISphere 的联合主席——一个服务全球 GIS 学生与研究者的志愿者社区。我从 2022 年的校园合伙人做起,后来负责 GISource 部门,发布过 50 多篇 GIS 留学申请博客(阅读量超 5 万),也为社区搭建了数据平台和 LLM 分析系统。此外,我对量化金融保持着长期兴趣。\n\n我也带着相机旅行。网站的「镜头之下」收录了我在世界各地拍的照片——每一个点,都是一个留在记忆里的地方。\n\n这个网站本身也是我的作品:从设计到上线全程 vibe coding,与 AI 协作完成。你正在对话的这个机器人,同样是它的一部分。",
    "faq": "# FAQ\n\n## 你目前在做什么工作?\n\n我在上海意鹰信息技术(Ebest Mobile)担任 AI 产品经理,2026 年 4 月至今,负责把 LLM 能力落进企业产品。\n\n## 你在找新机会吗?对什么样的机会感兴趣?\n\n目前全职在岗。对 AI 产品方向的交流与有意思的机会保持开放,欢迎邮件聊聊。\n\n## 可以怎么联系你?\n\n邮箱 sigao.li@outlook.com,或 LinkedIn(linkedin.com/in/sigao-li)。网站首页还有 GitHub 与 ResearchGate 链接。\n\n## 你的技术栈/核心能力是什么?\n\n我的能力组合是「AI 产品 × 数据工程 × 空间分析」:产品侧熟悉 Prompt Engineering、Function Calling、Agent Workflow;工程侧常用 Python、SQL、FastAPI,做过 ETL 管线与数据库设计;底子是六年 GIS 与空间分析训练,加上商业分析硕士。比起单项技能,我更擅长把这三者组合起来,从问题定义到上线完整交付。\n\n## 为什么从 GIS 转向 AI/产品?\n\n与其说转型,不如说延长线。地理分析教我用空间和关系理解问题,商业分析教我用数据回答商业问题;LLM 出现后,产品成了让这些能力真正被用上的方式。从地图到模型再到产品,每一步都是自然的下一步。\n\n## 你做过最有代表性的项目是什么?\n\n网站作品页有五个完整案例,最有代表性的三个:GISphere LLM 分析(多模态 LLM 管线,能读网页、PDF 和截图,把全球 GIS 学术机会结构化)、AI 财务分析门户(与 IBM 合作,上传年报即生成 SWOT/MOST/PESTLE 与情绪分析)、GISphere 数据平台(为全球志愿者组织打造的端到端数据产品)。每个案例都有「挑战—方案—影响」的完整叙述,欢迎追问细节。\n\n## 接受合作/咨询吗?\n\n与 AI 产品、数据分析、GIS 相关的合作与交流都欢迎,请邮件说明来意,我会回复。\n\n## 这个网站是怎么做的?\n\nAstro + Tailwind + GSAP 构建,托管在 GitHub Pages。网站页脚写着他的署名:「从设计到上线,全程 vibe coding。」(与 AI 协作)——注意这句只出现在网站上,他的简历里并没有写。页面里的生成式动效——作品页的等高线、简历页的河流时间线、首页的粒子场——都是为这个网站定制的。\n\n## 「镜头之下」的照片是你拍的吗?\n\n是,全部由我本人拍摄,版权保留。数量与足迹以「摄影足迹」统计为准(自动同步网站数据)——地图上每个点都可以点开看。\n\n## 你会记住我吗?会保存我的对话吗?\n\n本猫的记性只待在你自己的浏览器里喵——记得「你来过」、你更爱逛哪个版块(作品/简历/摄影,只记版块,不记你看了哪页),还有(要是你告诉过本猫)你的称呼,都存在你这台设备上,不在主人的服务器上;换个浏览器或清了缓存,本猫就重新不认识你了。聊天时本猫会顺手把「你更爱逛的版块」这一条(至多两个版块名)带给 AI,好把话聊到你关心的点子上。你打的字会发给一个第三方 AI 帮本猫想回答(顺便判断能不能给你指个相关页面),除此之外本猫不保存你的对话,网站的访问统计也是匿名的。想让本猫彻底忘掉你,点聊天框角落的「忘记我」就行,更细的说明在网站的「隐私说明」页喵。",
    "guidelines": "# 回答规范\n\n## 语气与风格\n\n- 专业但亲切,像 Sigao 本人在轻松场合的说话方式;不油腻、不营销腔、不堆砌成就\n- 跟随访客语言:中文提问用中文答,英文提问用英文答\n- 回答简短优先(2-5 句),访客追问再展开\n\n## 遇到不知道的问题\n\n- 知识范围里没有的信息,直说不了解,并引导访客邮件联系本人\n- 不编造、不推测 Sigao 的观点或经历;不确定就说不确定\n\n## 引导策略\n\n- 问经历细节 → 指路简历页(/cv)\n- 问项目细节 → 指路对应案例页(/work)\n- 想深入合作或长聊 → 引导邮件 sigao.li@outlook.com\n- 与 Sigao 无关的话题(代写代码、闲聊天气等)→ 友好地拉回主题,不硬拒",
    "boundaries": "# 边界清单\n\n## 不进知识包(构建期硬排除)\n\n- 精确家庭住址与家坐标(沿用照片隐私既有规则)\n- 手机号(邮箱和 LinkedIn 公开,电话不公开)\n- 证件号、生日等身份信息\n- 薪资信息(历史与期望)\n\n## 机器人拒答/绕开的话题\n\n- 评价具体的前雇主、前同事(可正面描述经历,不做负面评价)\n- 政治、宗教等立场话题\n- 替 Sigao 做承诺(报价、答应合作、约定时间)→ 一律引导邮件联系本人",
    "extra": "# 补充层(站外信息)\n\n- 常驻上海,时区 UTC+8。\n- GISphere 联合主席是志愿者角色,与本职工作相互独立。\n- 中英双语工作语言;在加拿大和英国共生活学习了约七年。"
  },
  "zh": {
    "cv": {
      "current": {
        "id": "ebest",
        "title": "AI产品经理",
        "org": "上海意鹰信息技术有限公司",
        "location": "中国上海",
        "start": "2026-04",
        "end": "present",
        "bullets": [],
        "url": "https://www.ebestmobile.cn/"
      },
      "experience": [
        {
          "id": "heywhale",
          "title": "咨询项目经理",
          "org": "上海和今信息科技有限公司",
          "location": "中国上海",
          "start": "2026-01",
          "end": "2026-04",
          "bullets": [
            "设计了一个基于 Web 的门户，利用 NLP 对年报、推文及回复进行情感分析",
            "执行了 MOST、SWOT 和 PESTLE 分析，并可视化呈现结果"
          ],
          "url": "https://www.heywhale.com/"
        },
        {
          "id": "piinpoint",
          "title": "地理空间数据分析师",
          "org": "PiinPoint",
          "location": "加拿大基奇纳",
          "start": "2023-01",
          "end": "2023-04",
          "bullets": [
            "实施 k-NN 客户细分与城市化指数进行市场筛选，优化零售网络达 15%",
            "将 ML 集成到 GIS 工作流中，并重构企业数据库架构，提升运营效率 30%"
          ],
          "url": "https://www.piinpoint.com/"
        },
        {
          "id": "iqiyi",
          "title": "商业分析师",
          "org": "上海爱奇艺文化传媒有限公司",
          "location": "中国上海",
          "start": "2021-06",
          "end": "2021-08",
          "bullets": [
            "对影视 IP 转化为线下业态及大型商业综合体趋势进行市场研究",
            "制定数据驱动的选址策略，并向公司高管汇报"
          ],
          "url": "https://www.iqiyi.com/"
        },
        {
          "id": "miotech",
          "title": "数据分析师",
          "org": "颖投信息科技 (上海) 有限公司",
          "location": "中国上海",
          "start": "2021-04",
          "end": "2021-06",
          "bullets": [
            "通过网络爬虫从政府和税务机构抓取经济数据进行 ESG 研究，将数据收集时间缩短 50%"
          ],
          "url": "https://www.miotech.com/zh-CN"
        },
        {
          "id": "sohu",
          "title": "产品分析师",
          "org": "北京搜狐互联网信息服务有限公司",
          "location": "中国北京",
          "start": "2021-02",
          "end": "2021-04",
          "bullets": [
            "进行竞品研究与用户行为诊断分析，提升用户参与度和留存率 20%"
          ],
          "url": "https://www.sohu.com/"
        }
      ],
      "research": [
        {
          "id": "ra-sustain",
          "title": "研究助理 — 可持续发展研究",
          "org": "布里斯托大学",
          "location": "英国布里斯托",
          "start": "2025-06",
          "end": "2025-08",
          "bullets": [
            "构建了一个基于 LLM 的系统，用于识别网页结构、抓取 CSR/ESG 报告并提取元数据，汇入可持续发展数据集"
          ],
          "url": "https://www.bristol.ac.uk/"
        },
        {
          "id": "ra-ai-or",
          "title": "研究助理 — AI与运筹学",
          "org": "布里斯托大学",
          "location": "英国布里斯托",
          "start": "2025-05",
          "end": "2025-10",
          "bullets": [
            "使用视觉语言模型提取环境线索，并构建数据集，用于微调 LLM 进行事故轨迹模拟"
          ],
          "url": "https://www.bristol.ac.uk/"
        },
        {
          "id": "ra-transport",
          "title": "研究助理 — 交通大数据",
          "org": "布里斯托大学",
          "location": "英国布里斯托",
          "start": "2024-12",
          "end": "2025-07",
          "bullets": [
            "利用 Graph Transformer + 贝叶斯优化（PyEPO）进行需求预测，为行程调度提供决策支持"
          ],
          "url": "https://www.bristol.ac.uk/"
        },
        {
          "id": "ra-nlp",
          "title": "研究助理 — 自然语言处理",
          "org": "布里斯托大学",
          "location": "英国布里斯托",
          "start": "2024-07",
          "end": "2025-04",
          "bullets": [
            "利用评分、评论和支出数据，评估英格兰政党控制权变更对养老院质量的影响"
          ],
          "url": "https://www.bristol.ac.uk/"
        },
        {
          "id": "ra-cal",
          "title": "研究助理 - 认知老化",
          "org": "多伦多都会大学认知老化实验室",
          "location": "加拿大多伦多",
          "start": "2024-03",
          "end": "2024-09",
          "bullets": [
            "为老龄化研究项目进行问卷预处理、翻译及研讨会协调工作"
          ],
          "url": "https://psychlabs.torontomu.ca/cal/"
        },
        {
          "id": "ra-health",
          "title": "研究助理 — 健康地理学",
          "org": "瑞尔森大学",
          "location": "加拿大多伦多",
          "start": "2022-09",
          "end": "2022-12",
          "bullets": [
            "对商店可达性与消费者健康进行回归分析；利用网络分析和地理编码进行消费者轨迹建模"
          ],
          "url": "https://www.torontomu.ca/"
        },
        {
          "id": "ra-gis",
          "title": "研究助理 — GIS",
          "org": "瑞尔森大学",
          "location": "加拿大多伦多",
          "start": "2022-07",
          "end": "2022-08",
          "bullets": [
            "构建 Python ETL 管道，聚合多伦多 GTA 选举结果，以考察选举多样性与包容性"
          ],
          "url": "https://www.torontomu.ca/"
        }
      ],
      "education": [
        {
          "id": "bristol-msc",
          "title": "商业分析硕士",
          "org": "布里斯托大学",
          "location": "英国布里斯托",
          "start": "2024-09",
          "end": "2025-11",
          "bullets": [
            "研究论文：理解与预测人类移动中的规律性、多样性与适应性"
          ],
          "url": "https://www.bristol.ac.uk/"
        },
        {
          "id": "tmu-msa",
          "title": "空间分析硕士",
          "org": "多伦多都会大学",
          "location": "加拿大多伦多",
          "start": "2022-09",
          "end": "2023-10",
          "bullets": [
            "研究论文：从地理视角分析劳氏公司在加拿大的失败"
          ],
          "url": "https://www.torontomu.ca/"
        },
        {
          "id": "ryerson-ba",
          "title": "地理分析荣誉学士，辅修经济学",
          "org": "瑞尔森大学",
          "location": "加拿大多伦多",
          "start": "2018-09",
          "end": "2022-06",
          "bullets": [
            "研究论文：多伦多都会区人口分布对少数族裔零售选址的影响"
          ],
          "url": "https://www.torontomu.ca/"
        }
      ],
      "volunteering": [
        {
          "id": "gisphere-cochair",
          "title": "联合主席",
          "org": "GISphere",
          "location": "全球",
          "start": "2026-05",
          "end": "present",
          "bullets": [
            "跨职能开发定制 LLM 聊天机器人，用于信息收集和技术支持",
            "领导 4 人团队设计基于 Azure 的 ETL 管道（Google Sheets → MySQL），将任务时间缩短 80%",
            "发布 50 余篇关于 GIS 项目申请的博客，阅读量超 5 万"
          ],
          "url": "https://gisphere.info/"
        },
        {
          "id": "gisource",
          "title": "部门负责人 - GISource",
          "org": "GISphere",
          "location": "全球",
          "start": "2024-04",
          "end": "2026-04",
          "bullets": [
            "跨职能开发定制 LLM 聊天机器人，用于信息收集和技术支持",
            "领导 4 人团队设计基于 Azure 的 ETL 管道（Google Sheets → MySQL），将任务时间缩短 80%",
            "发布 50 余篇关于 GIS 项目申请的博客，阅读量超 5 万"
          ],
          "url": "https://gisphere.info/"
        },
        {
          "id": "gisphere-campus",
          "title": "校园合伙人",
          "org": "GISphere",
          "location": "全球",
          "start": "2022-05",
          "end": "2024-03",
          "bullets": [
            "与 Esri 中国合作；组织“GIS 公开课周”（观看量超 1 万）；联合制作《GISphere 留学大数据白皮书（2023）》"
          ],
          "url": "https://gisphere.info/"
        },
        {
          "id": "cssa",
          "title": "学术助理",
          "org": "布里斯托中国学生学者联谊会",
          "location": "英国布里斯托",
          "start": "2024-07",
          "end": "2025-04",
          "bullets": [
            "策划学术资源，撰写关于申请时间线的微信公众号文章"
          ],
          "url": "https://www.bristolsu.org.uk/groups/cssa-chinese-students-scholars-association-b151"
        },
        {
          "id": "geoscene",
          "title": "校园大使",
          "org": "易智瑞信息技术有限公司",
          "location": "中国北京",
          "start": "2021-08",
          "end": "2022-08",
          "bullets": [
            "开展国际宣传活动，覆盖 300 余所大学和 1 万余名学生；资料下载量提升 15%"
          ],
          "url": "https://www.geoscene.cn/"
        }
      ],
      "awards": [
        {
          "year": "2024",
          "title": "Think Big 研究生奖学金（£6,500）"
        },
        {
          "year": "2023",
          "title": "加拿大制图协会制图竞赛会长奖"
        },
        {
          "year": "2023",
          "title": "研究生发展奖（C$700）"
        },
        {
          "year": "2022",
          "title": "文科研究生空间基金（C$5,000）"
        },
        {
          "year": "2021",
          "title": "全球高校 Python 量化模拟投资大赛 — 个人决赛第六名"
        },
        {
          "year": "2020",
          "title": "GLO-BUS 商业战略模拟 — 全球前 50 强"
        },
        {
          "year": "2018",
          "title": "保证可再生奖学金（C$500）"
        }
      ],
      "skills": [
        {
          "label": "AI 技术栈",
          "items": [
            "Prompt Engineering",
            "Function Calling",
            "Agent Workflow",
            "LangChain",
            "Autogen",
            "Redis",
            "Qdrant",
            "FastAPI"
          ]
        },
        {
          "label": "分析工具",
          "items": [
            "SQL",
            "Python",
            "A/B Testing",
            "Claude Code",
            "Codex",
            "Cursor",
            "Dify",
            "AI Studio",
            "NotebookLM",
            "Looker Studio"
          ]
        },
        {
          "label": "数据与工程",
          "items": [
            "Python",
            "R",
            "SQL",
            "NoSQL",
            "JavaScript",
            "Databricks",
            "Hadoop",
            "AWS",
            "Git",
            "Linux"
          ]
        },
        {
          "label": "机器学习框架",
          "items": [
            "PyTorch",
            "TensorFlow",
            "Keras",
            "scikit-learn",
            "NLTK",
            "NetworkX",
            "PuLP",
            "PySpark",
            "GeoPandas",
            "ArcPy"
          ]
        },
        {
          "label": "工具",
          "items": [
            "Tableau",
            "Power BI",
            "Google Analytics",
            "Esri Suite",
            "QGIS",
            "AutoCAD"
          ]
        }
      ],
      "certifications": [
        {
          "title": "人工智能",
          "org": "多伦多大学"
        },
        {
          "title": "数据科学",
          "org": "滑铁卢大学"
        },
        {
          "title": "PMP",
          "org": "在读"
        },
        {
          "title": "驾驶证",
          "org": "C1"
        },
        {
          "title": "CPR/AED",
          "org": "红十字会"
        }
      ]
    },
    "cases": [
      {
        "slug": "email-agent",
        "title": "智能邮件助手",
        "tagline": "一个住在飞书里的 LLM 邮件副驾驶——摘要、翻译、回复与记忆，尽在一张交互卡片中。",
        "year": "2026",
        "role": "personal",
        "repoUrl": "https://github.com/SigaoLi/INTELLIGENT_EMAIL_AGENT",
        "metrics": [
          {
            "label": "并行处理缩短 LLM 等待时间",
            "value": "~50%"
          },
          {
            "label": "步骤统一于一张累积式卡片中",
            "value": "3"
          },
          {
            "label": "持续学习的记忆维度",
            "value": "3"
          }
        ],
        "body": "## 挑战\n\n跨语言工作意味着每封邮件都要付出双倍时间：一次用来阅读，一次用来回复。现有的邮件客户端既不提供摘要，也没有上下文翻译，更不会记住谁写了什么——而在收件箱、翻译工具和聊天软件之间来回切换，一天要打断工作流几十次。\n\n目标：在不离开飞书的前提下完成\"阅读—起草—发送\"的完整闭环，让 AI 随时间推移记住往来对象和偏好——并且绝不未经人工审核就发送任何内容。\n\n## 方案\n\n**一张卡片，而非十条通知。** 助手通过 IMAP 以增量追踪的方式监控收件箱，将每封新邮件推送到一张飞书交互卡片中。阅读、生成回复和审核/发送作为三个步骤在同一张卡片内完成——已完成的步骤会自动折叠，因此繁忙的邮件往来绝不会淹没聊天窗口。\n\n**并行 LLM 流水线。** 摘要（中文摘要）和全文翻译作为并行的 LangChain 任务调用 Qwen-plus，与顺序调用相比，感知等待时间大约缩短了一半。慢速操作以异步方式派发，确保卡片始终即时响应。\n\n**持续积累的记忆。** 基于 mem0 和 ChromaDB 向量存储构建，助手维护三个记忆维度：联系人画像（他们是谁、如何写作）、跨邮件上下文（这个邮件线程究竟在讨论什么）以及用户偏好（语气、落款、决策）。每一次交互都会优化下一份草稿。\n\n**不事雕琢的正确性。** 完整维护 `In-Reply-To`/`References` 邮件头，确保线程在每个客户端中都保持完整；HTML 提取回退方案处理 Outlook 风格的纯 HTML 正文；APScheduler 驱动轮询；SQLite 跨重启追踪状态。\n\n## 影响\n\n这个助手将一项多工具、多语言的繁琐事务，转变为一个只需三次点击的审核流程——发送前始终有人工在回路中把关。作为一款产品，它展示了应用 AI 工艺的完整技术栈：延迟工程、平台约束下的交互设计，以及一套让系统在第四周明显优于第一周的记忆架构。\n\n**技术栈：** Python · LangChain · Qwen-plus · mem0 + ChromaDB · 飞书 WebSocket · IMAP/SMTP · APScheduler · SQLite"
      },
      {
        "slug": "gisphere-llm",
        "title": "GISphere LLM 分析",
        "tagline": "一个多模态 LLM 系统，能读懂网页、PDF 和截图，将全球 GIS 学术机会结构化。",
        "year": "2026",
        "role": "lead",
        "org": "GISphere (GIS-Info)",
        "repoUrl": "https://github.com/GIS-Info/GISPHERE_LLM_Analysis",
        "metrics": [
          {
            "label": "统一输入源类型数",
            "value": "5"
          },
          {
            "label": "分阶段 LLM 分析流水线",
            "value": "3"
          },
          {
            "label": "失效 API key 最短自动冷却时间（分钟）",
            "value": "30"
          }
        ],
        "body": "## 挑战\n\nGISphere 的志愿者需要追踪各种学术机会——博士招生、教职空缺、资助申请——它们散落在大学网页、PDF 传单、微信截图和共享表格里。把这些混乱信息变成干净、结构化的数据库，意味着每周要花数小时人工阅读和复制粘贴，数据质量完全取决于录入者是谁。\n\n作为主设计师和开发者，我的目标是让这条流水线能读懂志愿者扔给它的任何东西。\n\n## 方案\n\n**万物皆可读。** 系统能摄入五种来源——网页、PDF、截图、本地 Excel 和 Google Sheets。网页提取用 Playwright 做动态渲染，配合 trafilatura 提取干净文本；文档则依次经过 PyMuPDF → pdfplumber → Tesseract OCR → 视觉语言模型这一链条处理，哪怕是一张扫描版传单，最终也能变成结构化文本。\n\n**三阶段分析。** 提取出的内容会经过一个分阶段的 LLM 流水线：先识别出机会本身，再将其归类到 GIS 各子学科（自然地理、人文地理、城市、GIS、RS、GNSS），最后将结构化信息——截止日期、资助情况、联系方式——直接填入团队表格。\n\n**为不可靠的基础设施而设计。** 模型链网关能在 GPT、Gemini 和 Claude 之间自动降级切换；返回 401/403 的 API key 会进入 30 分钟的断路器冷却期；部分成功的行会保留已完成字段，而不是整行失败；批处理任务能从断点处恢复运行。数据入库前，还会通过 DuckDuckGo/Bing 进行搜索验证，交叉核对信息。\n\n## 影响\n\n该项目以 MIT 许可证在 GIS-Info 组织下发布，用一条有人监督的流水线取代了最枯燥的志愿者工作——人不再负责转录，而是负责核验。它是更宏大的 GISphere 数据平台的智能层，也是一个生产级 LLM 工程的实践样本：将优雅降级、多模态回退和故障隔离作为第一等的设计需求。\n\n**技术栈：** Python · 多模型网关（GPT / Gemini / Claude） · Playwright · trafilatura · PyMuPDF · Tesseract OCR · VLM · Google Sheets API"
      },
      {
        "slug": "gisphere-platform",
        "title": "GISphere 数据平台",
        "tagline": "从自动化管线到 BI 仪表盘与团队 KPI——为一个全球志愿者组织打造的端到端数据产品。",
        "year": "2025–2026",
        "role": "lead",
        "org": "GISphere",
        "repoUrl": "https://github.com/SigaoLi/GISPHERE_GOOGLE_SHEET",
        "metrics": [
          {
            "label": "ETL 管线减少的任务耗时",
            "value": "80%"
          },
          {
            "label": "50+ 篇已发布博客的累计阅读量",
            "value": "50k+"
          },
          {
            "label": "仪表盘中的可视化维度",
            "value": "10+"
          }
        ],
        "body": "## 挑战\n\nGISphere 为全球受众整理 GIS 研究生项目和就业市场信息，完全由志愿者运营。整个运作建立在电子表格之上：手动录入数据、手动发布微信公众号、对团队正在记录的就业市场缺乏全局视野，也无法判断团队自身的健康状况。\n\n作为 GISource 负责人，我主导了数据基础设施的建设——三个系统共同构成一个产品。\n\n## 方法\n\n**数据采集与发布自动化。** 一条 Python 管线将 Google Sheets 同步至 MySQL，通过 80/10/10 优先级算法筛选内容，自动检测新院校，校验必填字段，生成适配微信的发布内容并发送通知——具备 Gmail→QQ 邮箱自动降级与磁盘故障日志机制。采用模块化的 8 组件架构实现跨平台构建。\n\n**市场分析仪表盘。** 一个 Streamlit + Plotly 仪表盘读取合并后的 MySQL 与 Google Sheets 数据，通过 10+ 个可视化维度呈现全球 GIS 学术就业市场——时间序列、热力图、地图、Sankey 流向图、雷达图——并支持交互式多窗口切片。\n\n**团队 KPI 系统。** 第三层通过复合键（URL + 截止日期）将人工标注的表格数据与数据库匹配，以合理的规则处理模糊截止日期（如\"即将截止\"→30 天）来计算前置时间指标，并呈现成员贡献排名、每日趋势和地理覆盖情况。\n\n## 影响\n\n基于 Azure 的 ETL 管线由四人团队采用敏捷方法构建，将日常任务耗时削减了 80%。编辑产出达到 50+ 篇已发布博客，累计阅读量超过 50k。比各个部分更重要的是，整体展现了产品思维：一个数据模型同时服务于运营、分析和管理——对于一个靠志愿者工时运转的组织而言，这就是苦差与使命之间的区别。\n\n**技术栈：** Python · MySQL · Google Sheets/Docs API · Streamlit · Plotly · Pandas · Azure · APScheduler"
      },
      {
        "slug": "csr-scraper",
        "title": "ESG 报告智能系统",
        "tagline": "一套本地 LLM 驱动的抓取系统，用于发现、验证并分析企业可持续发展报告——在消费级硬件上完全离线运行。",
        "year": "2025",
        "role": "personal",
        "org": "University of Bristol (Research Assistant)",
        "repoUrl": "https://github.com/SigaoLi/UB_RA_CSR",
        "metrics": [
          {
            "label": "覆盖 49 家公司的目标报告数",
            "value": "150"
          },
          {
            "label": "通过严格模糊匹配实现的公司匹配精度",
            "value": "99%+"
          },
          {
            "label": "在可信平台快速通道上的速度提升",
            "value": "80–90%"
          }
        ],
        "body": "## 挑战\n\n可持续发展研究需要数十家上市公司跨越十年（2015–2024）的 CSR/ESG 报告——但这些报告隐藏在频繁改版的投资者关系网站、Cookie 弹窗、相似的公司名称和扫描版 PDF 之后。人工收集无法规模化；而简单的爬虫则会自信满满地抓取到错误公司的报告。\n\n作为布里斯托大学可持续发展项目的研究助理，我构建了这套系统，它还有一个额外的约束：完全本地运行，零 API 成本。\n\n## 方案\n\n**两个本地模型，分工协作。** 通过 Ollama 提供服务：DeepSeek-R1 8B 负责文本推理和搜索查询生成；Qwen2.5-VL 3B 负责视觉处理。扫描件或图片密集型 PDF 会进入纯视觉管线——页面被渲染成图像并由 VLM 读取——从而打破了经典的扫描文档处理僵局。\n\n**信任，但要验证——精度达 99%。** 公司身份由 LLM 驱动的模糊匹配进行验证，该匹配经过调优，要求极高的严格度（需达到 99%+ 的相似度），将错配率降至 0.1% 以下。来自可信报告平台的 PDF 会跳过完整的过滤链，在常规路径上实现了 80–90% 的速度提升。\n\n**一个像研究员一样工作的爬虫。** Playwright 驱动的导航能处理 Cookie 横幅，并跟踪最多五跳的多层级页面；GPU 加速的查重（CuPy）保持语料库的洁净；每一个决策都被记录下来以供审计。\n\n## 影响\n\n这套系统将原本需要耗费研究生大量时间的问题——定位并验证 49 家公司的 150 份报告——工业化成了一个在单张消费级 GPU 上运行的、受监督的批处理流程。在方法论上，它证明了精心的工程设计能让小型本地模型完成值得信赖的数据收集工作，而这通常是丢给昂贵的托管 API 去做的。\n\n**技术栈：** Python · Ollama · DeepSeek-R1 8B · Qwen2.5-VL 3B · Playwright · PyMuPDF · CuPy"
      },
      {
        "slug": "ibm-finance",
        "title": "AI 财务分析门户",
        "tagline": "上传年报，即可生成 SWOT、MOST 与 PESTLE 分析及情绪洞察——与 IBM 联合打造。",
        "year": "2025",
        "role": "team",
        "org": "IBM × 布里斯托大学",
        "repoUrl": "https://github.com/SigaoLi/UB_CP_IBM_Finance_Portal",
        "metrics": [
          {
            "label": "自动化战略框架",
            "value": "3"
          },
          {
            "label": "分析模式——引导式与自定义",
            "value": "2"
          },
          {
            "label": "行业合作伙伴",
            "value": "IBM"
          }
        ],
        "body": "## 挑战\n\n咨询顾问在接手任何项目时，头几天都花在从年报中提取战略信号上——每家公司动辄数百页，还要对每一家可比公司重复同样的流程。这个与 IBM 合作的咨询项目提出的问题是：一个网页门户能在多大程度上自动化这一轮初步分析，同时不丢失分析的结构性？\n\n## 方法\n\n**交付框架，而非仅仅摘要。** 门户接收 PDF 年报，生成结构化的 SWOT、MOST 和 PESTLE 分析——也就是咨询顾问实际会起草的那些交付物——同时提供贯穿整份报告、相关推文与回复的情绪分析，以及用于快速浏览的词云可视化。\n\n**两种模式，服务两类用户。** 引导模式运行标准分析流程，用于快速完成第一轮梳理；自定义模式则接收详细的用户需求，并据此制定实施策略。集成的对话界面引导用户完成上传与分析全过程。\n\n**务实的全栈构建。** Python NLP 后端搭配 JavaScript/HTML 前端，由一支商业分析咨询团队在四个月的项目周期内，按照 IBM 的需求简报设计并交付。\n\n## 影响\n\n该门户将年报的第一轮战略解读从数天压缩至数分钟，同时将输出保持在分析师实际使用的框架内——使其成为一件复核与精炼的工具，而非一个黑箱。作为一次行业合作，它既是 NLP 工程实践，也是咨询级别的需求转译。\n\n**技术栈：** Python · NLP/情绪分析管线 · JavaScript · HTML/CSS · Flask 风格网页门户"
      }
    ],
    "research": [
      {
        "title": "商业地理学",
        "summary": "人口结构、市场需求与空间竞争如何塑造零售战略——从多伦多族裔零售到劳氏退出加拿大市场。",
        "projects": [
          {
            "title": "电动汽车充电基础设施规划的多目标优化",
            "abstract": "结合多准则决策分析、混合整数规划与多目标优化方法，以布里斯托为案例进行电动汽车充电站选址研究。敏感性分析展示了在不同用户行为和交通假设下的适应性，为构建可持续、以用户为中心的充电网络提供了可复制的框架。",
            "github": "https://github.com/SigaoLi/UB_MA_Public_Facility_Site_Selection"
          },
          {
            "title": "从地理学视角分析劳氏在加拿大的失败",
            "abstract": "通过经营策略、财务表现和门店份额审视劳氏退出加拿大市场的原因，运用优化后的需求公式与Huff模型模拟市场需求。相对于家得宝，劳氏大部分门店位于低需求、高竞争的位置——为国际零售商的市场进入与运营提供了洞见。"
          },
          {
            "title": "多伦多大都市区人口分布对族裔零售选址的影响",
            "abstract": "利用人口普查和超市位置数据，考察多伦多大都市区华人超市的分布，通过缓冲区与泰森多边形估算销售潜力与商圈范围。研究发现：华人超市呈现郊区化趋势，选址以社区为锚点，并出现向服务南亚裔人群转变的新动向。"
          },
          {
            "title": "中小学扩招可行性分析",
            "abstract": "为布鲁斯-格雷天主教区教育局的法语沉浸式课程扩展项目进行评估，综合运用网络分析、Huff引力模型、多准则评价及选址-分配模型，提出了一套涵盖学区调整、设施升级和新校址的分层战略规划。"
          }
        ]
      },
      {
        "title": "犯罪分析",
        "summary": "面向循证警务的热点制图与凶杀模式分析——含一项 CCA 会长奖获奖作品。",
        "projects": [
          {
            "title": "多伦多市热点警务",
            "abstract": "地图海报展示热力图如何缩短警力响应时间、通过识别高犯罪区域提升巡逻效率，并通过定位事故多发地段为交通基础设施改善提供依据。获加拿大制图协会制图竞赛会长奖。",
            "link": "https://cca-acc.org/2023-cca-presidents-prize-winner-hotspot-policing-for-the-city-of-toronto.html"
          },
          {
            "title": "多伦多市凶杀类型分析",
            "abstract": "通过三项假设、文献综述与事件数据集的多图层制图，探索多伦多的凶杀模式——识别与凶杀率相关的因素，以支持预防与资源配置。"
          }
        ]
      },
      {
        "title": "环境监测",
        "summary": "基于卫星影像的深度学习，用于灾害检测与响应。",
        "projects": [
          {
            "title": "基于深度神经网络的森林火灾监测",
            "abstract": "一个从卫星影像中检测森林火灾的卷积神经网络，在验证集和测试集上展现出高准确率与低损失——证明了 AI 在灾害管理与响应中的潜力。",
            "github": "https://github.com/SigaoLi/UT_DL_Forest_Fire_Monitoring"
          }
        ]
      },
      {
        "title": "公共卫生",
        "summary": "健康结果的空间分析——从安大略省健康数据中的可塑面积单元问题，到英格兰地方政治如何影响养老院质量。",
        "projects": [
          {
            "title": "英格兰政党控制权变更对养老院质量的影响",
            "abstract": "对 116 个英格兰地方当局（2016–2023）的面板研究，将 CQC 评级与地方政府财政及社会经济数据整合进贝叶斯并行过程潜增长模型。长期一党控制未显示出显著效果；政党更替——尤其是保守党向工党的转变——稳健地提升了质量，且不受支出中介影响，挑战了既有的支出-质量传导路径假设。"
          },
          {
            "title": "安大略省健康中心数据聚合中的可塑面积单元问题",
            "abstract": "对南安大略省 366 个社区的多变量回归与预测分析显示，显著预测因子因地理聚合层级而异——高聚合层级构建的模型无法预测低聚合层级，对健康服务规划具有启示意义。"
          }
        ]
      },
      {
        "title": "量化金融",
        "summary": "从股价预测到实时欺诈检测——运用 Transformer、遗传算法与流式机器学习进行金融预测。",
        "projects": [
          {
            "title": "混合专家模型用于股价预测",
            "abstract": "运用 ARIMA、LSTM 与 MoE 模型预测苹果公司股价，基于准确率与财务回报进行评估。结合 ARIMA 与 LSTM 的 MoE 模型提供了最佳回报与稳定性，并为不同投资者类型量身定制了建议。",
            "github": "https://github.com/SigaoLi/UB_DA_Stock_Price_Prediction"
          },
          {
            "title": "遗传算法用于投资组合策略优化",
            "abstract": "将 Transformer 时序分析与遗传算法的搜索效率相结合，在提升预测准确率的同时，显著降低了计算开销。",
            "github": "https://github.com/SigaoLi/UT_AI_Portfolio_Strategy_Optimization"
          },
          {
            "title": "实时信用卡欺诈检测",
            "abstract": "一个实时欺诈预测框架：完整的 ML 管道开发，以及通过 Spark Streaming 和 Kafka 进行的流式交易处理。",
            "github": "https://github.com/SigaoLi/UW_BD_Credit_Card_Fraud_Detection"
          }
        ]
      },
      {
        "title": "网络分析",
        "summary": "通过多模态情感分析与动态主题建模，将非结构化的用户生成内容转化为战略性商业洞察。",
        "projects": [
          {
            "title": "健身、反馈与未来：通过社交媒体分析理解 PureGym 用户",
            "abstract": "将基于 BERT 的多模态情感分析与动态主题建模应用于 PureGym 的 Google Maps 评论。尽管 70% 的评论是正面的，但满意度随门店老化而下降；员工与停车位是好评的主要驱动因素，卫生问题则是投诉主因——这些运营洞察是传统评分所掩盖的。",
            "github": "https://github.com/SigaoLi/UB_SM_PUREGYM"
          }
        ]
      }
    ],
    "photos": {
      "totalPhotos": 76,
      "countries": [
        {
          "name": "中国",
          "photos": 16,
          "cities": [
            "成都",
            "峨眉山",
            "桂林",
            "大连",
            "南京",
            "敦煌",
            "茶卡",
            "秦皇岛",
            "北戴河",
            "承德"
          ],
          "descriptions": [
            "成都杜甫草堂，竹影环抱的池畔草亭",
            "峨眉山，云海之上的寺庙屋脊",
            "成都，古祠庭院的黄昏光线",
            "成都大熊猫基地，抱着竹子大快朵颐的大熊猫",
            "桂林，喀斯特山影环绕的湖面与红帆木船",
            "大连星海湾大桥的日落",
            "南京牛首山佛顶宫的彩绘长廊",
            "南京牛首山佛顶宫，经文墙尽头的鎏金佛像",
            "敦煌鸣沙山月牙泉，沙丘环抱的绿洲",
            "敦煌莫高窟九层楼",
            "青海茶卡盐湖的暮色",
            "秦皇岛开埠时期的欧式老建筑",
            "秦皇岛 1899 年开埠地老站台",
            "北戴河，渤海上的朦胧日出",
            "承德避暑山庄，柳岸拱桥",
            "承德磬锤峰"
          ]
        },
        {
          "name": "日本",
          "photos": 7,
          "cities": [
            "东京",
            "京都",
            "河口湖"
          ],
          "descriptions": [
            "东京猫头鹰咖啡馆里栖息的灰林鸮",
            "京都三十三间堂的悠长木构大殿",
            "京都清水寺的朱红仁王门与三重塔",
            "京都，黄昏时分穿透云层的霞光",
            "岚山竹林，仰望竹梢间的天光",
            "金阁寺与镜湖池",
            "河口湖芦苇丛外的富士山"
          ]
        },
        {
          "name": "加拿大",
          "photos": 16,
          "cities": [
            "多伦多",
            "阿尔冈昆公园",
            "班夫",
            "路易斯湖",
            "托伯莫里",
            "尼亚加拉瀑布城",
            "魁北克市",
            "贾斯珀"
          ],
          "descriptions": [
            "多伦多港，冬日落日下的货轮与吊臂",
            "黎明粉霞下的多伦多天际线",
            "安大略湖的秋日波光",
            "阿冈昆省立公园，静水与浮木",
            "班夫硫磺山顶望落基山雪原",
            "俯瞰雪中的班夫小镇",
            "冰封路易斯湖上的滑冰人",
            "班夫，新雪间蜿蜒的山涧",
            "落基山间的冰瀑",
            "托伯莫里布鲁斯半岛的清澈浅滩",
            "多伦多，五月晴空下的樱花",
            "尼亚加拉马蹄瀑布与雾中游船",
            "魁北克蒙莫朗西瀑布与悬索桥",
            "路易斯湖的暮色",
            "贾斯珀玛琳湖的精灵岛",
            "多伦多北郊的绯红日落"
          ]
        },
        {
          "name": "美国",
          "photos": 16,
          "cities": [
            "纽约",
            "费城",
            "华盛顿",
            "迈阿密",
            "旧金山",
            "圣马特奥",
            "蒙特雷",
            "圣巴巴拉",
            "洛杉矶",
            "圣莫尼卡"
          ],
          "descriptions": [
            "纽约时代广场，黄昏时分的霓虹",
            "纽约港上的自由女神像",
            "费城艺术博物馆前的华盛顿纪念喷泉雕像",
            "夕照中的林肯纪念堂",
            "暮色中亮起的二战纪念碑",
            "纽约河滨公园望哈德逊河",
            "布鲁克林大桥步道与曼哈顿天际线",
            "清晨哈德逊河上驶过的邮轮",
            "迈阿密海滩的清晨浪线",
            "马林岬角望金门大桥与旧金山",
            "加州一号公路旁的太平洋海岸",
            "蒙特雷十七里湾的孤柏",
            "圣巴巴拉法院钟楼俯瞰红瓦屋顶与海岸",
            "洛杉矶市中心的天使铁路缆车",
            "格里菲斯天文台山道远望好莱坞标志",
            "圣莫尼卡码头边的海滩"
          ]
        },
        {
          "name": "英国",
          "photos": 20,
          "cities": [
            "牛津",
            "剑桥",
            "伦敦",
            "布里斯托",
            "罗蒙湖",
            "因弗内斯",
            "南昆斯费里",
            "爱丁堡",
            "温莎",
            "科茨沃尔德",
            "贝尔法斯特",
            "卢尔沃斯",
            "格林尼治"
          ],
          "descriptions": [
            "牛津，学院蜜色石墙下的庭院",
            "剑桥康河上的数学桥",
            "剑桥国王学院礼拜堂",
            "伦敦肯辛顿花园的阿尔伯特纪念亭",
            "夏日云影下的白金汉宫",
            "布里斯托，横跨埃文峡谷的克利夫顿悬索桥",
            "晨雾中如镜的罗蒙湖",
            "因弗内斯，尼斯河畔的教堂尖顶",
            "南昆斯费里，列车驶过福斯铁路桥",
            "爱丁堡亚瑟王座山脚的春樱",
            "爱丁堡老学院圆顶街景",
            "泰晤士河上望伦敦塔桥与伦敦塔",
            "温莎城堡圣乔治礼拜堂",
            "科茨沃尔德乡村教堂的尖顶",
            "科茨沃尔德小路旁的茅草石屋",
            "贝尔法斯特市政厅与铜绿穹顶",
            "温莎城堡的卫兵换岗",
            "侏罗纪海岸卢尔沃斯，斯泰尔洞的褶皱岩层",
            "多塞特杜德尔门附近，Man O’War 湾的白垩悬崖",
            "格林尼治公园，旧皇家海军学院与金丝雀码头同框"
          ]
        },
        {
          "name": "巴哈马",
          "photos": 1,
          "cities": [
            "拿骚"
          ],
          "descriptions": [
            "拿骚太子乔治码头一字排开的邮轮"
          ]
        }
      ]
    }
  },
  "en": {
    "cv": {
      "current": {
        "id": "ebest",
        "title": "AI Product Manager",
        "org": "Ebest Mobile",
        "location": "Shanghai, China",
        "start": "2026-04",
        "end": "present",
        "bullets": [],
        "url": "https://www.ebestmobile.com/"
      },
      "experience": [
        {
          "id": "ibm-uob",
          "title": "Business Analytics Consultant",
          "org": "IBM & University of Bristol",
          "location": "Bristol, UK",
          "start": "2025-01",
          "end": "2025-04",
          "bullets": [
            "Designed a web-based portal using NLP for sentiment analysis on annual reports, tweets and responses",
            "Performed MOST, SWOT and PESTLE analysis with visualised results"
          ]
        },
        {
          "id": "piinpoint",
          "title": "Geospatial Data Analyst",
          "org": "PiinPoint",
          "location": "Kitchener, Canada",
          "start": "2023-01",
          "end": "2023-04",
          "bullets": [
            "Implemented k-NN customer segmentation and an urbanization index for market screening, optimizing the retail network by 15%",
            "Integrated ML into GIS workflows and revamped enterprise database schema, enhancing operational efficiency by 30%"
          ],
          "url": "https://www.piinpoint.com/"
        },
        {
          "id": "iqiyi",
          "title": "Business Analyst",
          "org": "iQIYI, Inc.",
          "location": "Shanghai, China",
          "start": "2021-06",
          "end": "2021-08",
          "bullets": [
            "Market research on transforming film & TV IP into offline ventures and trends in large brick-and-mortar complexes",
            "Devised data-driven site selection strategies presented to company executives"
          ],
          "url": "https://ir.iqiyi.com/"
        },
        {
          "id": "miotech",
          "title": "Data Analyst",
          "org": "MioTech",
          "location": "Shanghai, China",
          "start": "2021-04",
          "end": "2021-06",
          "bullets": [
            "ESG research via web scraping of economic data from government and revenue agencies, reducing data collection time by 50%"
          ],
          "url": "https://www.miotech.com/en-US"
        },
        {
          "id": "sohu",
          "title": "Product Analyst",
          "org": "Sohu.com Limited",
          "location": "Beijing, China",
          "start": "2021-02",
          "end": "2021-04",
          "bullets": [
            "Competitive research and diagnostic analysis of user behaviour, boosting user engagement and retention by 20%"
          ],
          "url": "https://investors.sohu.com/"
        }
      ],
      "research": [
        {
          "id": "ra-sustain",
          "title": "Research Assistant — Sustainable Development Research",
          "org": "University of Bristol",
          "start": "2025-06",
          "end": "present",
          "bullets": [
            "Built an LLM-based system to identify webpage structures, scrape CSR/ESG reports and extract metadata into a sustainability dataset"
          ],
          "url": "https://www.bristol.ac.uk/"
        },
        {
          "id": "ra-ai-or",
          "title": "Research Assistant — AI & Operations Research",
          "org": "University of Bristol",
          "start": "2025-05",
          "end": "present",
          "bullets": [
            "Used Vision Language Models to extract environmental cues and build a dataset for fine-tuning LLMs on accident trajectory simulation"
          ],
          "url": "https://www.bristol.ac.uk/"
        },
        {
          "id": "ra-transport",
          "title": "Research Assistant — Transportation Big Data",
          "org": "University of Bristol",
          "start": "2024-12",
          "end": "present",
          "bullets": [
            "Demand forecasting with Graph Transformer + Bayesian optimization (PyEPO) for trip dispatching decision support"
          ],
          "url": "https://www.bristol.ac.uk/"
        },
        {
          "id": "ra-nlp",
          "title": "Research Assistant — Natural Language Processing",
          "org": "University of Bristol",
          "start": "2024-07",
          "end": "2025-04",
          "bullets": [
            "Evaluated the impact of party control changes in England on care home quality using ratings, reviews and expenditure data"
          ],
          "url": "https://www.bristol.ac.uk/"
        },
        {
          "id": "ra-cal",
          "title": "Research Assistant",
          "org": "Cognitive Aging Lab, Toronto Metropolitan University",
          "start": "2024-03",
          "end": "2024-09",
          "bullets": [
            "Questionnaire preprocessing, translation and symposium coordination for aging research projects"
          ],
          "url": "https://psychlabs.torontomu.ca/cal/"
        },
        {
          "id": "ra-health",
          "title": "Research Assistant — Health Geography",
          "org": "Ryerson University",
          "start": "2022-09",
          "end": "2022-12",
          "bullets": [
            "Regression analysis of store accessibility vs. consumer health; consumer trajectory modelling with network analysis and geocoding"
          ],
          "url": "https://www.torontomu.ca/"
        },
        {
          "id": "ra-gis",
          "title": "Research Assistant — GIS",
          "org": "Ryerson University",
          "start": "2022-07",
          "end": "2022-08",
          "bullets": [
            "Python ETL pipeline aggregating Toronto GTA election results to examine electoral diversity and inclusion"
          ],
          "url": "https://www.torontomu.ca/"
        }
      ],
      "education": [
        {
          "id": "bristol-msc",
          "title": "MSc in Business Analytics",
          "org": "University of Bristol",
          "location": "Bristol, UK",
          "start": "2024-09",
          "end": "2025-11",
          "bullets": [
            "Research paper: Understanding and Predicting Regularity, Diversity, and Adaptability in Human Mobility"
          ],
          "url": "https://www.bristol.ac.uk/"
        },
        {
          "id": "tmu-msa",
          "title": "MSA in Spatial Analysis",
          "org": "Toronto Metropolitan University",
          "location": "Toronto, Canada",
          "start": "2022-09",
          "end": "2023-10",
          "bullets": [
            "Research paper: Analyzing Lowe's Failure in Canada from a Geographical Perspective"
          ],
          "url": "https://www.torontomu.ca/"
        },
        {
          "id": "ryerson-ba",
          "title": "BA (Hons) in Geographic Analysis, Minor in Economics",
          "org": "Ryerson University (now Toronto Metropolitan University)",
          "location": "Toronto, Canada",
          "start": "2018-09",
          "end": "2022-06",
          "bullets": [
            "Research paper: The Impact of Population Distribution in the Toronto CMA on Ethnic Retail Location"
          ],
          "url": "https://www.torontomu.ca/"
        }
      ],
      "volunteering": [
        {
          "id": "gisource",
          "title": "Director, GISource",
          "org": "GISphere",
          "start": "2024-04",
          "end": "present",
          "bullets": [
            "Cross-functional development of a custom LLM chatbot for information collection and technical support",
            "Led a team of 4 designing an Azure-based ETL pipeline (Google Sheets → MySQL), reducing task time by 80%",
            "Published 50+ blogs on GIS program applications with 50k+ reads"
          ],
          "url": "https://gisphere.info/"
        },
        {
          "id": "gisphere-campus",
          "title": "Campus Partner",
          "org": "GISphere",
          "start": "2022-05",
          "end": "2024-03",
          "bullets": [
            "Partnered with Esri China; organized 'GIS Open Course Week' (10k+ views); co-produced the GISphere Study Abroad Big Data White Paper (2023)"
          ],
          "url": "https://gisphere.info/"
        },
        {
          "id": "cssa",
          "title": "Academic Assistant",
          "org": "Chinese Students and Scholars Association Bristol",
          "start": "2024-07",
          "end": "present",
          "bullets": [
            "Academic resource curation and WeChat articles on admission timelines"
          ],
          "url": "https://www.bristolsu.org.uk/groups/cssa-chinese-students-scholars-association-b151"
        },
        {
          "id": "geoscene",
          "title": "Campus Ambassador",
          "org": "GeoScene Information Technology",
          "start": "2021-08",
          "end": "2022-08",
          "bullets": [
            "International campaigns reaching 300+ universities and 10k+ students; +15% material downloads"
          ],
          "url": "https://www.geoscene.cn/"
        }
      ],
      "awards": [
        {
          "year": "2024",
          "title": "Think Big Postgraduate Scholarship (£6,500)"
        },
        {
          "year": "2023",
          "title": "President's Prize, Canadian Cartographic Association Mapping Competition"
        },
        {
          "year": "2023",
          "title": "Graduate Development Award (C$700)"
        },
        {
          "year": "2022",
          "title": "Arts Grad Funding Spatial (C$5,000)"
        },
        {
          "year": "2021",
          "title": "Global University Python Quantitative Simulation Investment Competition — Individual Final 6th"
        },
        {
          "year": "2020",
          "title": "GLO-BUS Business Strategy Simulation — Global Top 50"
        },
        {
          "year": "2018",
          "title": "Guaranteed and Renewable Scholarship (C$500)"
        }
      ],
      "skills": [
        {
          "label": "Data & engineering",
          "items": [
            "Python",
            "R",
            "SQL",
            "NoSQL",
            "JavaScript",
            "Databricks",
            "Hadoop",
            "AWS",
            "Git",
            "Linux"
          ]
        },
        {
          "label": "ML frameworks",
          "items": [
            "PyTorch",
            "TensorFlow",
            "Keras",
            "scikit-learn",
            "NLTK",
            "NetworkX",
            "PuLP",
            "PySpark",
            "GeoPandas",
            "ArcPy"
          ]
        },
        {
          "label": "Tools",
          "items": [
            "Tableau",
            "Power BI",
            "Google Analytics",
            "Esri Suite",
            "QGIS",
            "AutoCAD"
          ]
        }
      ],
      "certifications": [
        {
          "title": "Artificial Intelligence",
          "org": "University of Toronto"
        },
        {
          "title": "Data Science",
          "org": "University of Waterloo"
        }
      ]
    },
    "cases": [
      {
        "slug": "email-agent",
        "title": "Intelligent Email Agent",
        "tagline": "An LLM email copilot living in Feishu — summaries, translations, replies and memory in one interactive card.",
        "year": "2026",
        "role": "personal",
        "repoUrl": "https://github.com/SigaoLi/INTELLIGENT_EMAIL_AGENT",
        "metrics": [
          {
            "label": "LLM wait time cut by parallel processing",
            "value": "~50%"
          },
          {
            "label": "steps unified in one accumulating card",
            "value": "3"
          },
          {
            "label": "memory dimensions that keep learning",
            "value": "3"
          }
        ],
        "body": "## Challenge\n\nWorking across languages means every email costs twice: once to read it, once to answer it. Existing clients offer no summarization, no contextual translation, and no memory of who writes what — and switching between a mailbox, a translator and a chat tool breaks flow dozens of times a day.\n\nThe goal: handle the entire read–draft–send loop without leaving Feishu, with an AI that remembers correspondents and preferences over time — and never sends anything without human review.\n\n## Approach\n\n**One card, not ten notifications.** The agent monitors the inbox over IMAP with incremental tracking, and pushes each new email into a single Feishu interactive card. Reading, reply generation and review/send happen as three steps inside the same card — completed steps fold away automatically, so a busy thread never floods the chat.\n\n**Parallel LLM pipeline.** Summarization (Chinese digest) and full translation run as parallel LangChain tasks against Qwen-plus, cutting perceived wait time roughly in half compared to sequential calls. Slow operations are dispatched asynchronously so the card always responds instantly.\n\n**Memory that compounds.** Built on mem0 with a ChromaDB vector store, the agent maintains three memory dimensions: contact profiles (who they are, how they write), cross-email context (what this thread is really about), and user preferences (tone, sign-offs, decisions). Every interaction refines the next draft.\n\n**Unglamorous correctness.** Full `In-Reply-To`/`References` header maintenance keeps threads intact in every client; an HTML-extraction fallback handles Outlook-style HTML-only bodies; APScheduler drives polling; SQLite tracks state across restarts.\n\n## Impact\n\nThe agent turns a multi-tool, multi-language chore into a three-tap review flow — with a human always in the loop before send. As a product, it demonstrates the full stack of applied-AI craft: latency engineering, interaction design under platform constraints, and a memory architecture that makes the system measurably better in week four than in week one.\n\n**Stack:** Python · LangChain · Qwen-plus · mem0 + ChromaDB · Feishu WebSocket · IMAP/SMTP · APScheduler · SQLite"
      },
      {
        "slug": "gisphere-llm",
        "title": "GISphere LLM Analysis",
        "tagline": "A multimodal LLM system that reads webpages, PDFs and screenshots to structure the world's GIS academic opportunities.",
        "year": "2026",
        "role": "lead",
        "org": "GISphere (GIS-Info)",
        "repoUrl": "https://github.com/GIS-Info/GISPHERE_LLM_Analysis",
        "metrics": [
          {
            "label": "input source types unified",
            "value": "5"
          },
          {
            "label": "stage LLM analysis pipeline",
            "value": "3"
          },
          {
            "label": "min auto-cooldown on failing API keys",
            "value": "30"
          }
        ],
        "body": "## Challenge\n\nGISphere volunteers track academic opportunities — PhD openings, faculty positions, funding calls — scattered across university pages, PDF flyers, WeChat screenshots and shared spreadsheets. Turning that chaos into a clean, structured database meant hours of manual reading and copy-pasting per week, with quality depending entirely on who did the typing.\n\nAs the lead designer and developer, I set out to make the pipeline read anything a volunteer could throw at it.\n\n## Approach\n\n**Read anything.** The system ingests five source types — webpages, PDFs, screenshots, local Excel and Google Sheets. Web extraction uses Playwright for dynamic rendering with trafilatura for clean text; documents fall through a chain of PyMuPDF → pdfplumber → Tesseract OCR → vision-language model, so even a scanned flyer ends as structured text.\n\n**Three-stage analysis.** Extracted content passes through a staged LLM pipeline that identifies the opportunity, classifies it across GIS sub-disciplines (Physical Geo, Human Geo, Urban, GIS, RS, GNSS), and fills the structured schema — deadlines, funding, contacts — directly into the team's sheet.\n\n**Engineered for unreliable infrastructure.** A model-chain gateway falls back across GPT, Gemini and Claude; API keys that return 401/403 enter a 30-minute circuit-breaker cooldown; partially successful rows keep their completed fields instead of failing whole; and batch runs resume from where they stopped. Search verification cross-checks claims via DuckDuckGo/Bing before data lands.\n\n## Impact\n\nReleased as an MIT-licensed project under the GIS-Info organization, the system replaces the most tedious volunteer workflow with a supervised pipeline — humans verify instead of transcribe. It is the intelligence layer of the broader GISphere data platform, and a working study in production LLM engineering: graceful degradation, multimodal fallbacks, and failure isolation as first-class design requirements.\n\n**Stack:** Python · multi-model gateway (GPT / Gemini / Claude) · Playwright · trafilatura · PyMuPDF · Tesseract OCR · VLM · Google Sheets API"
      },
      {
        "slug": "gisphere-platform",
        "title": "GISphere Data Platform",
        "tagline": "From automation pipeline to BI dashboards and team KPIs — an end-to-end data product for a global volunteer organization.",
        "year": "2025–2026",
        "role": "lead",
        "org": "GISphere",
        "repoUrl": "https://github.com/SigaoLi/GISPHERE_GOOGLE_SHEET",
        "metrics": [
          {
            "label": "task time reduced by the ETL pipeline",
            "value": "80%"
          },
          {
            "label": "reads across 50+ published blogs",
            "value": "50k+"
          },
          {
            "label": "visualization dimensions in the dashboard",
            "value": "10+"
          }
        ],
        "body": "## Challenge\n\nGISphere curates GIS graduate-program and job-market information for a worldwide audience, run entirely by volunteers. The operation lived in spreadsheets: manual data entry, manual WeChat publishing, no view of the job market the team was documenting, and no way to see whether the team itself was healthy.\n\nAs Director of GISource, I led the build-out of the data infrastructure — three systems that together form one product.\n\n## Approach\n\n**Ingestion & publishing automation.** A Python pipeline syncs Google Sheets into MySQL, selects content via an 80/10/10 priority algorithm, auto-detects new universities, validates required fields, generates WeChat-ready content and sends notifications — with Gmail→QQmail automatic fallback and failure logs on disk. Built cross-platform with a modular 8-component architecture.\n\n**Market analytics dashboard.** A Streamlit + Plotly dashboard reads the merged MySQL + Google Sheets data and exposes the global GIS academic job market across 10+ visualization dimensions — time series, heatmaps, maps, Sankey flows, radar charts — with interactive multi-window slicing.\n\n**Team KPI system.** A third layer matches human-annotated sheet data to the database via composite keys (URL + deadline), computes lead-time metrics with sensible rules for fuzzy deadlines (\"Soon\" → 30 days), and surfaces member contribution rankings, daily trends and geographic coverage.\n\n## Impact\n\nThe Azure-based ETL pipeline, built with a team of four using agile methods, cut routine task time by 80%. Editorial output reached 50+ published blogs with 50k+ cumulative reads. More than the parts, the whole demonstrates product thinking: one data model serving operations, analytics and management — for an organization that runs on volunteer hours, the difference between a chore and a mission.\n\n**Stack:** Python · MySQL · Google Sheets/Docs API · Streamlit · Plotly · Pandas · Azure · APScheduler"
      },
      {
        "slug": "csr-scraper",
        "title": "ESG Report Intelligence",
        "tagline": "A local-LLM scraping system that finds, validates and analyzes corporate sustainability reports — air-gapped, on consumer hardware.",
        "year": "2025",
        "role": "personal",
        "org": "University of Bristol (Research Assistant)",
        "repoUrl": "https://github.com/SigaoLi/UB_RA_CSR",
        "metrics": [
          {
            "label": "reports targeted across 49 companies",
            "value": "150"
          },
          {
            "label": "company-match precision via strict fuzzy matching",
            "value": "99%+"
          },
          {
            "label": "faster on trusted-platform fast paths",
            "value": "80–90%"
          }
        ],
        "body": "## Challenge\n\nSustainability research needs a decade of CSR/ESG reports (2015–2024) for dozens of public companies — but the reports hide behind redesigned investor-relations sites, cookie walls, lookalike company names and scanned PDFs. Manual collection doesn't scale; naive scraping collects the wrong company's reports with confidence.\n\nBuilt as a research assistant on the University of Bristol's sustainable development project, the system had one more constraint: run fully local, with zero API cost.\n\n## Approach\n\n**Two local models, divided labor.** Served via Ollama: DeepSeek-R1 8B handles text reasoning and search-query generation; Qwen2.5-VL 3B handles vision. Scanned or image-heavy PDFs route through a pure-vision pipeline — pages rendered to images and read by the VLM — breaking the classic scanned-document deadlock.\n\n**Trust, but verify — at 99%.** Company identity is validated by LLM-driven fuzzy matching tuned for extreme strictness (99%+ similarity required), pushing mismatch rates below 0.1%. PDFs from trusted report platforms skip the full filter chain, an 80–90% speedup on the common path.\n\n**A scraper that behaves like a researcher.** Playwright-driven navigation handles cookie banners and follows up to five hops of multi-level pages; GPU-accelerated deduplication (CuPy) keeps the corpus clean; every decision is logged for audit.\n\n## Impact\n\nThe system industrializes what was a graduate-student-hours problem — locating and validating 150 reports across 49 companies — into a supervised batch process on a single consumer GPU. Methodologically, it shows that careful engineering lets small local models do trustworthy data-collection work that's usually thrown at expensive hosted APIs.\n\n**Stack:** Python · Ollama · DeepSeek-R1 8B · Qwen2.5-VL 3B · Playwright · PyMuPDF · CuPy"
      },
      {
        "slug": "ibm-finance",
        "title": "AI Financial Analysis Portal",
        "tagline": "Upload an annual report, get SWOT, MOST and PESTLE analyses with sentiment — built with IBM.",
        "year": "2025",
        "role": "team",
        "org": "IBM × University of Bristol",
        "repoUrl": "https://github.com/SigaoLi/UB_CP_IBM_Finance_Portal",
        "metrics": [
          {
            "label": "strategy frameworks automated",
            "value": "3"
          },
          {
            "label": "analysis modes — guided & custom",
            "value": "2"
          },
          {
            "label": "industry partner",
            "value": "IBM"
          }
        ],
        "body": "## Challenge\n\nConsultants spend their first days on any engagement extracting strategy signals from annual reports — hundreds of pages per company, repeated across every comparable. The IBM-partnered consulting project asked: how much of that first-pass analysis can a web portal automate without losing analytical structure?\n\n## Approach\n\n**Frameworks, not just summaries.** The portal accepts a PDF annual report and generates structured SWOT, MOST and PESTLE analyses — the actual artifacts a consultant would draft — alongside sentiment analysis across the report, related tweets and responses, and word-cloud visualizations for fast scanning.\n\n**Two modes for two audiences.** A guided mode runs the standard analysis battery for a quick first pass; a custom mode takes detailed user requirements and develops implementation strategies against them. An integrated chat interface walks users through upload and analysis.\n\n**Pragmatic full-stack build.** Python NLP backend with a JavaScript/HTML front end, designed and delivered by a Business Analytics consulting team working to IBM's brief over a four-month engagement.\n\n## Impact\n\nThe portal compresses the first-pass strategic read of an annual report from days to minutes, while keeping outputs in the frameworks analysts actually use — making it a review-and-refine tool rather than a black box. As an industry collaboration, it was equal parts NLP engineering and consulting-grade requirement translation.\n\n**Stack:** Python · NLP/sentiment pipeline · JavaScript · HTML/CSS · Flask-style web portal"
      }
    ],
    "research": [
      {
        "title": "Business Geography",
        "summary": "How demographic factors, market demand and spatial competition shape retail strategy — from ethnic retail in Toronto to Lowe's exit from Canada.",
        "projects": [
          {
            "title": "Multi-Objective Optimization for EV Charging Infrastructure Planning",
            "abstract": "Combines multi-criteria decision analysis, mixed integer programming and multi-objective optimization for EV charging station siting, using Bristol as a case study. Sensitivity analysis demonstrates adaptability under different user-behaviour and traffic assumptions, yielding a replicable framework for sustainable, user-focused charging networks.",
            "github": "https://github.com/SigaoLi/UB_MA_Public_Facility_Site_Selection"
          },
          {
            "title": "Analyzing Lowe's Failure in Canada from a Geographical Perspective",
            "abstract": "Examines Lowe's Canadian exit through operating strategy, financial performance and store share, simulating market demand with an optimized demand formula and the Huff Model. Most Lowe's stores sat in low-demand, high-competition locations relative to The Home Depot — insights for international retailers on market entry and operations."
          },
          {
            "title": "The Impact of Population Distribution in the Toronto CMA on Ethnic Retail Location",
            "abstract": "Uses census and supermarket location data to examine Chinese supermarket distribution in the Toronto CMA, estimating sales potential and trade areas via buffer zones and Thiessen polygons. Findings: suburbanization of Chinese supermarkets, community-anchored siting, and an emerging shift toward serving South Asian populations."
          },
          {
            "title": "Feasibility Analysis of Enrollment Expansion in Primary and Secondary Schools",
            "abstract": "Evaluated French Immersion program expansion for the Bruce-Grey Catholic District School Board using network analysis, Huff gravity modelling, multi-criteria evaluation and location-allocation models — delivering a tiered strategic plan for boundary adjustments, facility upgrades and new school sites."
          }
        ]
      },
      {
        "title": "Crime Analysis",
        "summary": "Hotspot mapping and homicide pattern analysis for evidence-based policing — including a CCA President's Prize-winning map.",
        "projects": [
          {
            "title": "Hotspot Policing for the City of Toronto",
            "abstract": "Map poster showing how heat maps can reduce police response time, improve patrol efficiency by detecting high-crime areas, and inform transportation infrastructure improvements by locating accident-prone areas. Winner of the President's Prize at the Canadian Cartographic Association Mapping Competition.",
            "link": "https://cca-acc.org/2023-cca-presidents-prize-winner-hotspot-policing-for-the-city-of-toronto.html"
          },
          {
            "title": "Analysis of Homicides by Type in the City of Toronto",
            "abstract": "Explores homicide patterns in Toronto through three hypotheses, literature review and multi-layer mapping of incident datasets — identifying factors that correlate with homicide rates to support prevention and resource allocation."
          }
        ]
      },
      {
        "title": "Environmental Monitoring",
        "summary": "Deep learning on satellite imagery for disaster detection and response.",
        "projects": [
          {
            "title": "Forest Fire Monitoring Using Deep Neural Networks",
            "abstract": "A convolutional neural network that detects forest fires from satellite imagery, showing high accuracy and low loss on validation and test data — demonstrating AI's potential in disaster management and response.",
            "github": "https://github.com/SigaoLi/UT_DL_Forest_Fire_Monitoring"
          }
        ]
      },
      {
        "title": "Public Health",
        "summary": "Spatial analysis of health outcomes — from the MAUP in Ontario health data to how local politics shape care home quality in England.",
        "projects": [
          {
            "title": "The Impact of Changes in Party Political Control in England on Care Home Quality",
            "abstract": "Panel study of 116 English local authorities (2016–2023) integrating CQC ratings with local government financial and socioeconomic data in a Bayesian parallel-process latent growth model. Long-term partisan control shows no significant effect; political alternation — especially Conservative-to-Labour transitions — robustly improves quality, unmediated by expenditure, challenging the assumed spending–quality pathway."
          },
          {
            "title": "Modifiable Areal Unit Problem in Ontario Health Central Data Aggregation",
            "abstract": "Multivariate regression and predictive analysis across 366 southern-Ontario neighbourhoods shows significant predictors differ across geographic aggregation levels — models built at high aggregation fail to predict at low aggregation, with implications for health services planning."
          }
        ]
      },
      {
        "title": "Quantitative Finance",
        "summary": "Transformers, genetic algorithms and streaming ML for financial prediction — from stock prices to real-time fraud detection.",
        "projects": [
          {
            "title": "Mixture of Experts for Stock Price Prediction",
            "abstract": "Predicts Apple Inc. stock price with ARIMA, LSTM and MoE models, evaluated on accuracy and financial return. The MoE model combining ARIMA and LSTM delivers the best returns and stability, with tailored recommendations for different investor types.",
            "github": "https://github.com/SigaoLi/UB_DA_Stock_Price_Prediction"
          },
          {
            "title": "Genetic Algorithms for Portfolio Strategy Optimization",
            "abstract": "Combines Transformer time-series analysis with genetic algorithm search efficiency, significantly reducing computational overhead while improving prediction accuracy.",
            "github": "https://github.com/SigaoLi/UT_AI_Portfolio_Strategy_Optimization"
          },
          {
            "title": "Real-Time Credit Card Fraud Detection",
            "abstract": "A real-time fraud prediction framework: full ML pipeline development plus streamed transaction processing with Spark Streaming and Kafka.",
            "github": "https://github.com/SigaoLi/UW_BD_Credit_Card_Fraud_Detection"
          }
        ]
      },
      {
        "title": "Web Analytics",
        "summary": "Turning unstructured user-generated content into strategic business intelligence via multimodal sentiment analysis and dynamic topic modelling.",
        "projects": [
          {
            "title": "Fitness, Feedback, and the Future: Understanding PureGym Users Through Social Media Analytics",
            "abstract": "Applies BERT-based multimodal sentiment analysis and dynamic topic modelling to PureGym's Google Maps reviews. While 70% of reviews are positive, satisfaction declines as locations age; staff and parking drive praise, hygiene drives complaints — operational insights traditional ratings obscure.",
            "github": "https://github.com/SigaoLi/UB_SM_PUREGYM"
          }
        ]
      }
    ],
    "photos": {
      "totalPhotos": 76,
      "countries": [
        {
          "name": "China",
          "photos": 16,
          "cities": [
            "Chengdu",
            "Mount Emei",
            "Guilin",
            "Dalian",
            "Nanjing",
            "Dunhuang",
            "Chaka",
            "Qinhuangdao",
            "Beidaihe",
            "Chengde"
          ],
          "descriptions": [
            "Thatched pavilions over a quiet pond at Du Fu's Thatched Cottage, Chengdu",
            "Temple rooftops above a sea of clouds on Mount Emei",
            "Evening light in a historic temple courtyard, Chengdu",
            "A giant panda settled into its bamboo lunch, Chengdu Panda Base",
            "Red-sailed wooden boats on a karst-ringed lake, Guilin",
            "Sunset behind Xinghai Bay Bridge, Dalian",
            "Painted gallery leading into Usnisa Palace, Niushou Mountain, Nanjing",
            "A gilded Buddha beyond a wall of engraved sutras, Usnisa Palace, Nanjing",
            "Crescent Lake oasis amid the Mingsha dunes, Dunhuang",
            "The nine-storey pavilion of the Mogao Caves, Dunhuang",
            "Dusk over the salt flats of Chaka Salt Lake, Qinghai",
            "Colonial-era architecture from Qinhuangdao’s port-opening days",
            "The 1899 heritage railway stop at Qinhuangdao",
            "Hazy sunrise over the Bohai Sea at Beidaihe",
            "An arched bridge among willows at the Chengde Mountain Resort",
            "Sledgehammer Rock standing over Chengde"
          ]
        },
        {
          "name": "Japan",
          "photos": 7,
          "cities": [
            "Tokyo",
            "Kyoto",
            "Lake Kawaguchi"
          ],
          "descriptions": [
            "A tawny owl perched at a Tokyo owl café",
            "The long wooden hall of Sanjūsangen-dō, Kyoto",
            "The vermilion gate and pagoda of Kiyomizu-dera, Kyoto",
            "Sun rays breaking over the Kyoto basin at dusk",
            "Looking up through the Arashiyama bamboo grove",
            "Kinkaku-ji, the Golden Pavilion, across its mirror pond",
            "Mount Fuji over the reeds of Lake Kawaguchi"
          ]
        },
        {
          "name": "Canada",
          "photos": 16,
          "cities": [
            "Toronto",
            "Algonquin Park",
            "Banff",
            "Lake Louise",
            "Tobermory",
            "Niagara Falls",
            "Québec City",
            "Jasper"
          ],
          "descriptions": [
            "Winter sunset among ships at the Port of Toronto",
            "Dawn breaking over the Toronto skyline",
            "Autumn sun glitter on Lake Ontario",
            "Still water and driftwood, Algonquin Provincial Park",
            "Winter Rockies panorama from Sulphur Mountain, Banff",
            "Banff townsite from above, wrapped in snow",
            "Skaters far out on frozen Lake Louise",
            "A mountain creek threading through fresh snow, Banff",
            "A frozen waterfall in the Canadian Rockies",
            "Clear shallows at the Bruce Peninsula, Tobermory",
            "Cherry blossoms against a May sky, Toronto",
            "Horseshoe Falls and the mist boat, Niagara",
            "Montmorency Falls under a suspension bridge, Québec",
            "Evening calm at Lake Louise",
            "Spirit Island on Maligne Lake, Jasper",
            "Crimson sunset over north Toronto"
          ]
        },
        {
          "name": "United States",
          "photos": 16,
          "cities": [
            "New York",
            "Philadelphia",
            "Washington, D.C.",
            "Miami",
            "San Francisco",
            "San Mateo",
            "Monterey",
            "Santa Barbara",
            "Los Angeles",
            "Santa Monica"
          ],
          "descriptions": [
            "Times Square neon at dusk, New York",
            "The Statue of Liberty from New York Harbor",
            "The Washington Monument fountain before the Philadelphia Museum of Art",
            "The Lincoln Memorial in late-day light",
            "The World War II Memorial glowing at dusk",
            "The Hudson River from Riverside Park, New York",
            "On the Brooklyn Bridge promenade, Manhattan behind",
            "A cruise ship slips up the Hudson at dawn",
            "Morning surf at Miami Beach",
            "The Golden Gate Bridge from the Marin Headlands",
            "Pacific surf along the San Mateo coast",
            "The Lone Cypress on 17-Mile Drive, Monterey",
            "Red-tile rooftops of Santa Barbara from the courthouse tower",
            "Angels Flight, the tiny funicular in downtown LA",
            "The Hollywood Sign from the Griffith Observatory trails",
            "Santa Monica beach from the pier"
          ]
        },
        {
          "name": "United Kingdom",
          "photos": 20,
          "cities": [
            "Oxford",
            "Cambridge",
            "London",
            "Bristol",
            "Loch Lomond",
            "Inverness",
            "South Queensferry",
            "Edinburgh",
            "Windsor",
            "Cotswolds",
            "Belfast",
            "Lulworth",
            "Greenwich"
          ],
          "descriptions": [
            "Honey-stone quadrangle gardens at an Oxford college",
            "The Mathematical Bridge over the Cam, Cambridge",
            "King's College Chapel from the Backs, Cambridge",
            "The Albert Memorial, Kensington Gardens, London",
            "Buckingham Palace under summer clouds",
            "Clifton Suspension Bridge over the Avon Gorge, Bristol",
            "Loch Lomond, mirror-still under morning haze",
            "Church spires along the River Ness, Inverness",
            "A train crossing the Forth Bridge, South Queensferry",
            "Spring blossoms below Arthur’s Seat, Edinburgh",
            "The Old College dome from South Bridge, Edinburgh",
            "Tower Bridge and the Tower of London from the Thames",
            "St George's Chapel inside Windsor Castle",
            "A village church spire in the Cotswolds",
            "Thatched cottages along a Cotswolds lane",
            "Belfast City Hall and its copper domes",
            "The Changing of the Guard at Windsor Castle",
            "Folded limestone at Stair Hole, Lulworth, Jurassic Coast",
            "Man O'War Bay's chalk cliffs near Durdle Door, Dorset",
            "Greenwich Park: the Old Royal Naval College against Canary Wharf"
          ]
        },
        {
          "name": "Bahamas",
          "photos": 1,
          "cities": [
            "Nassau"
          ],
          "descriptions": [
            "Cruise ships lined up at Nassau's Prince George Wharf"
          ]
        }
      ]
    }
  }
}