How to extract data from html table in shell script?

html, html-parsing, regex, sed, shell

Solution

Go with `(g)awk`, it's capable :-), here is a solution, but please note: it's only working with the exact html table format you had posted.

 awk -F "</*td>|</*tr>" '/<\/*t[rd]>.*[A-Z][A-Z]/ {print $3, $5, $7 }' FILE

Here you can see it in action: https://ideone.com/zGfLe

Some explanation:

`-F` sets the input field separator to a regexp (any of `tr`'s or `td`'s opening or closing tag

then works only on lines that matches those tags AND at least two upercasse fields

then prints the needed fields.

HTH

Problem

I am trying to create a BASH script what would extract the data from HTML table. Below is the example of table from where I need to extract data: ``` <table border=1> <tr> <td><b>Component</b></td> <td><b>Status</b></td> <td><b>Time / Error</b></td> </tr> <tr><td>SAVE_DOCUMENT</td><td>OK</td><td>0.406 s</td></tr> <tr><td>GET_DOCUMENT</td><td>OK</td><td>0.332 s</td></tr> <tr><td>DVK_SEND</td><td>OK</td><td>0.001 s</td></tr> <tr><td>DVK_RECEIVE</td><td>OK</td><td>0.001 s</td></tr> <tr><td>GET_USER_INFO</td><td>OK</td><td>0.143 s</td></tr> <tr><td>NOTIFICATIONS</td><td>OK</td><td>0.001 s</td></tr> <tr><td>ERROR_LOG</td><td>OK</td><td>0.001 s</td></tr> <tr><td>SUMMARY_STATUS</td><td>OK</td><td>0.888 s</td></tr> </table> ``` And I want the BASH script to output it like so: ``` SAVE_DOCUMENT OK 0.475 s GET_DOCUMENT OK 0.345 s DVK_SEND OK 0.002 s DVK_RECEIVE OK 0.001 s GET_USER_INFO OK 4.465 s NOTIFICATIONS OK 0.001 s ERROR_LOG OK 0.002 s SUMMARY_STATUS OK 5.294 s ``` How to do it? So far I have tried using the sed, but I don't know how to use it quite well. The header of the table(Component, Status, Time/Error) I excluded with grep using `grep "<tr><td>`, so only lines starting with `<tr><td>` will be selected for next parsing (sed). This is what I used: `sed 's@<\([^<>][^<>]*\)>\([^<>]*\)</\1>@\2@g'` But then `<tr>` tags still remain and also it wont separate the strings. In other words the result of this script is: ``` <tr>SAVE_DOCUMENTOK0.406 s</tr> ``` The full command of the script I'm working on is: ``` cat $FILENAME | grep "<tr><td>" | sed 's@<\([^<>][^<>]*\)>\([^<>]*\)</\1>@\2@g' ```

Original source