How can I scrape a page with dynamic content (created by JavaScript) in Python?
javascript, python, web-scraping
Solution
EDIT Sept 2021: `phantomjs` isn't maintained any more, either
EDIT 30/Dec/2017: This answer appears in top results of Google searches, so I decided to update it. The old answer is still at the end.
dryscape isn't maintained anymore and the library dryscape developers recommend is Python 2 only. I have found using Selenium's python library with Phantom JS as a web driver fast enough and easy to get the work done.
Once you have installed Phantom JS, make sure the `phantomjs` binary is available in the current path:
phantomjs --version
# result:
2.1.1
#Example To give an example, I created a sample page with following HTML code. (link):
<!DOCTYPE html>
<html>
<head>
<meta charset="utf-8">
<title>Javascript scraping test</title>
</head>
<body>
<p id='intro-text'>No javascript support</p>
<script>
document.getElementById('intro-text').innerHTML = 'Yay! Supports javascript';
</script>
</body>
</html>
without javascript it says: `No javascript support` and with javascript: `Yay! Supports javascript`
#Scraping without JS support:
import requests
from bs4 import BeautifulSoup
response = requests.get(my_url)
soup = BeautifulSoup(response.text)
soup.find(id="intro-text")
# Result:
<p id="intro-text">No javascript support</p>
#Scraping with JS support:
from selenium import webdriver
driver = webdriver.PhantomJS()
driver.get(my_url)
p_element = driver.find_element_by_id(id_='intro-text')
print(p_element.text)
# result:
'Yay! Supports javascript'
You can also use Python library dryscrape to scrape javascript driven websites.
#Scraping with JS support:
import dryscrape
from bs4 import BeautifulSoup
session = dryscrape.Session()
session.visit(my_url)
response = session.body()
soup = BeautifulSoup(response)
soup.find(id="intro-text")
# Result:
<p id="intro-text">Yay! Supports javascript</p>
Problem
I'm trying to develop a simple web scraper. I want to extract plain text without HTML markup. My code works on plain (static) HTML, but not when content is generated by JavaScript embedded in the page. In particular, when I use `urllib2.urlopen(request)` to read the page content, it doesn't show anything that would be added by the JavaScript code, because that code isn't executed anywhere. Normally it would be run by the web browser, but that isn't a part of my program. How can I access this dynamic content from within my Python code? See also Can scrapy be used to scrape dynamic content from websites that are using AJAX? for answers specific to Scrapy.
Related problems
- How to get a Docker container's IP address from the host
- Can scrapy be used to scrape dynamic content from websites that are using AJAX?
- How to convert raw javascript object to a dictionary?
- How to Fix JSON Key Values without double-quotes?
- How to use Cors anywhere to reverse proxy and add CORS headers
- Launch javascript function from pyqt QWebEngineView
- Requests-html results in OSError: [Errno 8] Exec format error when calling html.render()