Text Extraction from HTML Java
html, html-content-extraction, java, screen-scraping, text-extraction
Solution
jsoup
Another html parser I really liked using was jsoup. You could get all the `<p>` elements in 2 lines of code.
Document doc = Jsoup.connect("http://en.wikipedia.org/").get();
Elements ps = doc.select("p");
Then write it out to a file in one more line
out.write(ps.text()); //it will append all of the p elements together in one long string
or if you want them on separate lines you can iterate through the elements and write them out separately.
Problem
I'm working on a program that downloads HTML pages and then selects some of the information and write it to another file. I want to extract the information which is intbetween the paragraph tags, but i can only get one line of the paragraph. My code is as follows; ``` FileReader fileReader = new FileReader(file); BufferedReader buffRd = new BufferedReader(fileReader); BufferedWriter out = new BufferedWriter(new FileWriter(newFile.txt)); String s; while ((s = br.readLine()) !=null) { if(s.contains("<p>")) { try { out.write(s); } catch (IOException e) { } } } ``` i was trying to add another while loop, which would tell the program to keep writing to file until the line contains the `</p>` tag, by saying; ``` while ((s = br.readLine()) !=null) { if(s.contains("<p>")) { while(!s.contains("</p>") { try { out.write(s); } catch (IOException e) { } } } } ``` But this doesn't work. Could someone please help.