// the find
Trinkle23897/learn2018-autodown
清华大学新版网络学堂课程自动下载脚本 / A python script to clone all files from learn.tsinghua.edu.cn
A Python script that logs into Tsinghua's new Web Learning portal (learn.tsinghua.edu.cn) and mirrors course announcements, slides, assignments with grading feedback, discussions and course info into local folders. It is for Tsinghua students who want an offline copy of a semester's material, and it is useless outside that portal.
- Login runs in a real Chrome window, so two-factor and SSO flows are handled by the browser. The script only extracts the resulting session, which avoids reimplementing a login protocol that the university changes.
- Incremental sync compares files and downloads only new or changed content. The --clear flag removes byte-identical duplicates across folders, so re-running after the first full pull is cheap.
- Filtering by semester, course, and ignore list, combined with --multi and a process count, makes partial syncs practical rather than all-or-nothing.
- Unattended runs are hard. When the session expires, the documented fix is deleting session.json and re-running, which reopens Chrome for a manual login. Cron jobs or headless servers need a desktop session with a browser.
- Scraping is brittle. The README's troubleshooting asks users to grep the source for an 'imcompatible os' comment and edit User-Agent handling by hand. The last push was 2025-09-19, so nothing in the repo shows it still works against the current portal.
- The tree shows no test directory, and the cross-platform claim ('完美运行于 Windows、macOS 和 Linux') is not backed by anything visible in the repo. Selenium and gmssl are unpinned in the README, so a dependency bump could break it without warning.