I have been following this tutorial: https://kb.objectrocket.com/postgresql/scrape-a-website-to-postgres-with-python-938
My app.py file looks like this (taken from the above tutorial):
from flask import Flask # needed for flask-dependent libraries below
from flask import render_template # to render the error page
from selenium import webdriver # to grab source from URL
from bs4 import BeautifulSoup # for searching through HTML
import psycopg2 # for database access
# set up Postgres database connection and cursor.
t_host = "localhost" # either "localhost", a domain name, or an IP address.
t_port = "5432" # default postgres port
t_dbname = "scrape"
t_user = "postgres"
t_pw = "********"
db_conn = psycopg2.connect(host=t_host, port=t_port, dbname=t_dbname, user=t_user, password=t_pw)
db_cursor = db_conn.cursor()
app = Flask(__name__)
def import_temp():
# set up your webdriver to use Chrome web browser
my_web_driver = webdriver.Chrome("/usr/local/bin/chromedriver")
# designate the URL we want to scrape
# NOTE: the long string of characters at the end of this URL below is a clue that
# maybe this page is so dynamic, like maybe refers to a specific web session and/or day/time,
# that we can't necessarily count on it to be the same more than one time.
# Which means... we may want to find another source for our data; one that is more
# dependable. That said, whatever URL you use, the methodology in this lesson stands.
t_url = "https://weather.com/weather/today/l/7ebb344012f0c5ff88820d763da89ed94306a86c770fda50c983bf01a0f55c0d"
# initiate scrape of website page data
my_web_driver.get("<a href='" t_url "'>" t_url "</a>")
# return entire page into "t_content"
t_content = my_web_driver.page_source
# use soup to make page content easily searchable
soup_in_bowl = BeautifulSoup(t_content)
# search for the UNIQUE span and class for the data we are looking for:
o_temp = soup_in_bowl.find('span', attrs={'class': 'deg-feels'})
# from the resulting object, "o_temp", get the text parameter and assign it to "n_temp"
n_temp = o_temp.text
# Build SQL for purpose of:
# saving the temperature data to a new row
s = ""
s = "INSERT INTO tbl_temperatures"
s = "("
s = "n_temp"
s = ") VALUES ("
s = "(%n_temp)"
s = ")"
# Trap errors for opening the file
db_cursor.execute(s, [n_temp, n_temp])
except psycopg2.Error as e:
t_msg = "Database error: " e "/n open() SQL: " s
return render_template("error_page.html", t_msg = t_msg)
# Success!
# Show a message to user.
t_msg = "Successful scrape!"
return render_template("progress.html", t_msg = t_msg)
# Clean up the cursor and connection objects
From looking at the Python Shell Error Log it looks like the URL is invalid:
FLASK_APP = app.py
FLASK_ENV = development
In folder /home/lloyd/PycharmProjects/flaskProject
/home/lloyd/PycharmProjects/flaskProject/venv/bin/python -m flask run
* Serving Flask app 'app.py' (lazy loading)
* Environment: development
* Debug mode: off
* Running on (Press CTRL C to quit)
[2021-12-28 18:13:05,988] ERROR in app: Exception on / [GET]
Traceback (most recent call last):
File "/home/lloyd/PycharmProjects/flaskProject/venv/lib/python3.6/site-packages/flask/app.py", line 2073, in wsgi_app
response = self.full_dispatch_request()
File "/home/lloyd/PycharmProjects/flaskProject/venv/lib/python3.6/site-packages/flask/app.py", line 1518, in full_dispatch_request
rv = self.handle_user_exception(e)
File "/home/lloyd/PycharmProjects/flaskProject/venv/lib/python3.6/site-packages/flask/app.py", line 1516, in full_dispatch_request
rv = self.dispatch_request()
File "/home/lloyd/PycharmProjects/flaskProject/venv/lib/python3.6/site-packages/flask/app.py", line 1502, in dispatch_request
return self.ensure_sync(self.view_functions[rule.endpoint])(**req.view_args)
File "/home/lloyd/PycharmProjects/flaskProject/app.py", line 33, in import_temp
my_web_driver.get("<a href='" t_url "'>" t_url "</a>")
File "/home/lloyd/PycharmProjects/flaskProject/venv/lib/python3.6/site-packages/selenium/webdriver/remote/webdriver.py", line 333, in get
self.execute(Command.GET, {'url': url})
File "/home/lloyd/PycharmProjects/flaskProject/venv/lib/python3.6/site-packages/selenium/webdriver/remote/webdriver.py", line 321, in execute
File "/home/lloyd/PycharmProjects/flaskProject/venv/lib/python3.6/site-packages/selenium/webdriver/remote/errorhandler.py", line 242, in check_response
raise exception_class(message, screen, stacktrace)
selenium.common.exceptions.WebDriverException: Message: unknown error: unhandled inspector error: {"code":-32000,"message":"Cannot navigate to invalid URL"}
(Session info: chrome=96.0.4664.110)
(Driver info: chromedriver=2.35.528139 (47ead77cb35ad2a9a83248b292151462a66cd881),platform=Linux 4.18.0-259.el8.x86_64 x86_64) - - [28/Dec/2021 18:13:05] "GET / HTTP/1.1" 500 -
However, I am able to visit the URL when I manually enter the address.
When I run the application a web console appears with the error msg:
Internal Server Error
The server encountered an internal error and was unable to complete your request. Either the server is overloaded or there is an error in the application.
Then a second web console appears with the text data:,
in the address bar of the web browser.
Any insight as to why this is happening would be greatly appreciated. I did previously post a similar question re a 404 ERROR here: Why am I receiving ERROR 404 - when attempting to use Python Flask
CodePudding user response:
All these errors:
[2021-12-28 18:13:05,988] ERROR in app: Exception on / [GET]
- selenium.common.exceptions.WebDriverException: Message: unknown error: unhandled inspector error: {"code":-32000,"message":"Cannot navigate to invalid URL"}
- web console appears with the text
are due to the the incompatibility between the version of the binaries you are using.
- You are using chrome=96.0.4664.45
- Release Notes of ChromeDriver v96.0 clearly mentions the following :
Supports Chrome version 96
- But you are using chromedriver=2.35.528139
- Release Notes of chromedriver=2.35.528139 clearly mentions the following :
Supports Chrome v62-64
So there is a clear mismatch between chromedriver=91.0 and the chrome=2.35.528139
Ensure that:
- ChromeDriver is updated to current ChromeDriver v96.0 level.
- Chrome is updated to current chrome=96.0.4664.45 (as per chrome=96.0.4664.45 release notes).