Your own weekly Python-news archive.
/build/project-hn-python-scraper with the starter and the rubric already in it.Fetch HN's Python tag feed, parse the front-page entries, dedupe by URL, and save the result to a JSON file. Run it weekly to build an archive of Python conversations you can search.
You'll never learn `requests`, `BeautifulSoup`, or file I/O from a lecture — you learn them by scraping something you actually want. This one takes about an hour and gives you a script you'll re-run.
Build against this exact list — it is the spec your published project is judged on.
Copy this into a new file — or open it in a workspace as scraper.py. The TODO parts are yours to fill.
"""
HN Python scraper.
Usage:
python scraper.py
Writes hn-python.json with a list of {title, url, points, discuss_url}.
"""
import json
from pathlib import Path
import requests
from bs4 import BeautifulSoup
URL = "https://news.ycombinator.com/from?site=python.org"
OUT = Path("hn-python.json")
def fetch() -> str:
# TODO: fetch URL with a real User-Agent header and return response.text.
# Retry once on failure. Raise for status codes 4xx/5xx.
# In the PyRun workspace the browser blocks news.ycombinator.com (it
# sends no CORS headers). Run this step with "python scraper.py" on your
# machine, or practise here against the browser-friendly JSON API:
# https://hn.algolia.com/api/v1/search?tags=story&query=python
raise NotImplementedError("step 1, fetch(): download URL and return the HTML")
def parse(html: str) -> list[dict]:
# TODO: BeautifulSoup(html, "html.parser") — extract each story row.
# Return list of {"title", "url", "points", "discuss_url"}.
raise NotImplementedError("step 2, parse(): turn the HTML into a list of story dicts")
def dedupe(items: list[dict]) -> list[dict]:
# TODO: keep first occurrence per URL.
raise NotImplementedError("step 3, dedupe(): keep the first story per URL")
def save(items: list[dict]) -> None:
OUT.write_text(json.dumps(items, indent=2, ensure_ascii=False))
def main() -> None:
html = fetch()
items = parse(html)
items = dedupe(items)
save(items)
print(f"Saved {len(items)} entries to {OUT}")
if __name__ == "__main__":
try:
main()
except NotImplementedError as todo:
# A fresh starter is SUPPOSED to stop here. Say which step is next
# instead of printing a traceback that looks like a bug.
print(f"Not built yet: {todo}")
print("Write that function, then press Run again. Each step you finish moves this message forward.")
scraper.py and a README.md holding all 6 rubric items as a checklist.pyrun.in/u/<handle>/w/project-hn-python-scraper.pyrun.in/u/<handle> portfolio page.Your workspace saves to this browser only. Sign in to keep it across devices and to publish it.
Push your solution to a public GitHub gist, repo, or Replit — paste the URL below. Optional: paste your `scraper.py` inline for the AI reviewer to give line-level feedback.
A published PyRun workspace URL (pyrun.in/u/<handle>/w/project-hn-python-scraper) works as the public URL too.