Wie vermeide ich HTML-Head-Tags in Jsoup? Parse

Question

Oct 03, 2014, 07:36 AM

Wie vermeide ich HTML-Head-Tags in Jsoup? Parse

Mit Jsoup versuche ich den angegebenen HTML-Inhalt zu analysieren. Nach Jsoup.parse () hängt die HTML-Ausgabe das Tag html, head und body an die Eingabe an. Ich möchte diese einfach ignorieren.

Sample Input:

<p><b>This <i>is</i></b> <i>my sentence</i> of text.</p>

Java Code:

import java.io.File;
import java.io.IOException;

import org.apache.commons.io.FileUtils;
import org.jsoup.Jsoup;
import org.jsoup.nodes.Document;
import org.jsoup.nodes.Element;
import org.jsoup.select.Elements;

public class HTMLParse {

    public static void main(String args[]) throws IOException {
        try{
            File input = new File("/ab.html");
            String html = FileUtils.readFileToString(input, null);

            Document doc = Jsoup.parseBodyFragment(html);
            doc.outputSettings().prettyPrint(false);
            System.out.println(doc.html());
        }
        catch(Exception e){
            e.printStackTrace();
        }
    }
}

Aktuelle Ausgabe:

<html><head></head><body><p><b>This <i>is</i></b> <i>my sentence</i> of text.</p>
    </body></html>

Erwartete Ausgabe

<p><b>This <i>is</i></b> <i>my sentence</i> of text.</p>

Bitte hilfe.

Zu kommentieren

Antworten auf die Frage(3)

Ihre Antwort auf die Frage

Top Fragen

0 die antwort

pandas plot dataframe barplot mit farben nach kategorie

0 die antwort

socket.gaierror: [Errno -2] Name oder Dienst nicht bekannt

0 die antwort

Selbstverknüpfung von Tabellen in Rails nicht möglich

0 die antwort

wie man 60 Matrizen in R schnell kombiniert

0 die antwort

So erhalten Sie eine öffentliche Variable (in einem Modul), um den Wert NICHT zwischen Benutzern zu teilen