scraping: download files from url
curl, python, r, selenium, web-scraping
Solution
Python version that use `BeautifulSoup`.
try:
# Python 3.x
from urllib.request import urlopen, urlretrieve, quote
from urllib.parse import urljoin
except ImportError:
# Python 2.x
from urllib import urlopen, urlretrieve, quote
from urlparse import urljoin
from bs4 import BeautifulSoup
url = 'http://oilandgas.ky.gov/Pages/ProductionReports.aspx'
u = urlopen(url)
try:
html = u.read().decode('utf-8')
finally:
u.close()
soup = BeautifulSoup(html)
for link in soup.select('div[webpartid] a'):
href = link.get('href')
if href.startswith('javascript:'):
continue
filename = href.rsplit('/', 1)[-1]
href = urljoin(url, quote(href))
try:
urlretrieve(href, filename)
except:
print('failed to download')
Problem
I want to automatically download files from this page. I tried many methods like: ``` download.file read.table GET ``` But without success. I am not asking for code , but I am asking for any hint/idea to deal with such situation.