Я шукаю швидкий спосіб отримати код відповіді HTTP із URL-адреси (тобто 200, 404 тощо). Я не впевнений, яку бібліотеку використовувати.
Я шукаю швидкий спосіб отримати код відповіді HTTP із URL-адреси (тобто 200, 404 тощо). Я не впевнений, яку бібліотеку використовувати.
Відповіді:
Оновіть за допомогою чудової бібліотеки запитів . Зверніть увагу, що ми використовуємо запит HEAD, що має відбуватися швидше, ніж повний запит GET або POST.
import requests
try:
r = requests.head("https://stackoverflow.com")
print(r.status_code)
# prints the int of the status code. Find more at httpstatusrappers.com :)
except requests.ConnectionError:
print("failed to connect")
requestsдає 403для вашого посилання, хоча все ще працює в браузері.
Ось рішення, яке використовує httplibзамість цього.
import httplib
def get_status_code(host, path="/"):
""" This function retreives the status code of a website by requesting
HEAD data from the host. This means that it only requests the headers.
If the host cannot be reached or something else goes wrong, it returns
None instead.
"""
try:
conn = httplib.HTTPConnection(host)
conn.request("HEAD", path)
return conn.getresponse().status
except StandardError:
return None
print get_status_code("stackoverflow.com") # prints 200
print get_status_code("stackoverflow.com", "/nonexistant") # prints 404
exceptблок хоча б для StandardErrorтого, щоб ви неправильно ловили такі речі KeyboardInterrupt.
curl -I http://www.amazon.com/.
Ви повинні використовувати urllib2, наприклад:
import urllib2
for url in ["http://entrian.com/", "http://entrian.com/does-not-exist/"]:
try:
connection = urllib2.urlopen(url)
print connection.getcode()
connection.close()
except urllib2.HTTPError, e:
print e.getcode()
# Prints:
# 200 [from the try block]
# 404 [from the except block]
http://entrian.com/на http://entrian.com/blogу своєму прикладі, отриманий результат 200 був би правильним, навіть якщо він передбачав переспрямування на http://entrian.com/blog/(зверніть увагу на кінцеву скісну риску).
Надалі для тих, хто використовує python3 і пізніші версії, ось ще один код для пошуку коду відповіді.
import urllib.request
def getResponseCode(url):
conn = urllib.request.urlopen(url)
return conn.getcode()
urllib2.HTTPErrorВиключення не містить getcode()метод. codeЗамість цього використовуйте атрибут.
Ось httplibрішення, яке поводиться як urllib2. Ви можете просто дати йому URL-адресу, і це просто працює. Не потрібно возитися з розподілом URL-адрес на ім’я хосту та шлях. Ця функція це вже робить.
import httplib
import socket
def get_link_status(url):
"""
Gets the HTTP status of the url or returns an error associated with it. Always returns a string.
"""
https=False
url=re.sub(r'(.*)#.*$',r'\1',url)
url=url.split('/',3)
if len(url) > 3:
path='/'+url[3]
else:
path='/'
if url[0] == 'http:':
port=80
elif url[0] == 'https:':
port=443
https=True
if ':' in url[2]:
host=url[2].split(':')[0]
port=url[2].split(':')[1]
else:
host=url[2]
try:
headers={'User-Agent':'Mozilla/5.0 (X11; Ubuntu; Linux x86_64; rv:26.0) Gecko/20100101 Firefox/26.0',
'Host':host
}
if https:
conn=httplib.HTTPSConnection(host=host,port=port,timeout=10)
else:
conn=httplib.HTTPConnection(host=host,port=port,timeout=10)
conn.request(method="HEAD",url=path,headers=headers)
response=str(conn.getresponse().status)
conn.close()
except socket.gaierror,e:
response="Socket Error (%d): %s" % (e[0],e[1])
except StandardError,e:
if hasattr(e,'getcode') and len(e.getcode()) > 0:
response=str(e.getcode())
if hasattr(e, 'message') and len(e.message) > 0:
response=str(e.message)
elif hasattr(e, 'msg') and len(e.msg) > 0:
response=str(e.msg)
elif type('') == type(e):
response=e
else:
response="Exception occurred without a good error message. Manually check the URL to see the status. If it is believed this URL is 100% good then file a issue for a potential bug."
return response