Remove All html tag except one tag
c#, regex
Solution
Regex is not great for parsing XML or HTML. Take a look at the HTML Agility Pack
HTML Agility Pack
Problem
I have some code to remove all html tag but I want to remove all html but except `</td>` and `</tr>` tags. How can this be done? ``` public string HtmlStrip( string input) { input = Regex.Replace(input, "<input>(.|\n)*?</input>", "*"); input = Regex.Replace(input, @"<xml>(.|\n)*?</xml>", "*"); // remove all <xml></xml> tags and anything inbetween. return Regex.Replace(input, @"<(.|\n)*?>", "*"); // remove any tags but not there content "<p>bob<span> johnson</span></p>" becomes "bob johnson" } ```