Python Projects & Skills / The web and automation
Lesson 20 of 28
Web scraping basics
BeautifulSoup reads HTML so you can pull out titles, links and tables. Check a site's terms and robots.txt, and don't hammer servers with requests.
Key points
- requests gets the page, BeautifulSoup parses it
- soup.find_all('a') gets links
- Be polite: add delays
from bs4 import BeautifulSoup
soup = BeautifulSoup(html, "html.parser")
for a in soup.find_all("a"):
print(a.get("href"))Quiz · +10 XP
What should you check before scraping a site?
Log in to save progress and earn XP.