Requests - get content-type/size without fetching the whole page/content
content-type, http-headers, python, python-requests
Solution
Yes.
You can use the `Session.head` method to create `HEAD` requests:
response = session.head(url, timeout=self.pageOpenTimeout, headers=customHeaders)
contentType = response.headers['content-type']
A `HEAD` request similar to `GET` request, except that the message body would not be sent.
Here is a quote from Wikipedia:
HEAD Asks for the response identical to the one that would correspond to a GET request, but without the response body. This is useful for retrieving meta-information written in response headers, without having to transport the entire content.
Problem
I have a simple website crawler, it works fine, but sometime it stuck because of large content such as ISO images, .exe files and other large stuff. Guessing content-type using file extension is probably not the best idea. Is it possible to get content-type and content length/size without fetching the whole content/page? Here is my code: ``` requests.adapters.DEFAULT_RETRIES = 2 url = url.decode('utf8', 'ignore') urlData = urlparse.urlparse(url) urlDomain = urlData.netloc session = requests.Session() customHeaders = {} if maxRedirects == None: session.max_redirects = self.maxRedirects else: session.max_redirects = maxRedirects self.currentUserAgent = self.userAgents[random.randrange(len(self.userAgents))] customHeaders['User-agent'] = self.currentUserAgent try: response = session.get(url, timeout=self.pageOpenTimeout, headers=customHeaders) currentUrl = response.url currentUrlData = urlparse.urlparse(currentUrl) currentUrlDomain = currentUrlData.netloc domainWWW = 'www.' + str(urlDomain) headers = response.headers contentType = str(headers['content-type']) except: logging.basicConfig(level=logging.DEBUG, filename=self.exceptionsFile) logging.exception("Get page exception:") response = None ```