Lazo Lab Sign up

Python Projects & Skills / The web and automation

Lesson 20 of 28

Web scraping basics

BeautifulSoup reads HTML so you can pull out titles, links and tables. Check a site's terms and robots.txt, and don't hammer servers with requests.

Key points

  • requests gets the page, BeautifulSoup parses it
  • soup.find_all('a') gets links
  • Be polite: add delays
from bs4 import BeautifulSoup
soup = BeautifulSoup(html, "html.parser")
for a in soup.find_all("a"):
    print(a.get("href"))
Watch a video on thisOpens YouTube search results for “Web scraping basics” in a new tab

Quiz · +10 XP

What should you check before scraping a site?

Log in to save progress and earn XP.