How to get unicode string when extract data in Python?

Question

I am trying to extract text from a Vietnamese website, which charset is in utf-8. However, the text I got is always in Ascii, and I can't find a way to convert them to unicode or get exactly the text on the website. As a result, I can't save them into file as expected.
I know this is the very popular problem with unicode in Python, but I still hope someone will help me to figure it out. Thanks.
My code:

import requests, re, io
import simplejson as json
from lxml import html, etree

base = "http://www.amthuc365.vn/cong-thuc/"
page = requests.get(base + "trang-" + str(1) + ".html")
pageTree = html.fromstring(page.text)

links = pageTree.xpath('//ul[contains(@class, "mt30")]/li/a/@href')
names = pageTree.xpath('//h3[@class="title"]/a/text()')
for name in names[:1]:
    print name
    # LÃ m bÃ¡nh oreo nhÃ¢n bÆ¡ Äáºu phá»ng thÆ¡m bÃ¹i

but what I need is "Làm bánh oreo nhân bơ đậu phộng thơm bùi"
Thanks.

score 2 · Accepted Answer · edited May 23 '17 at 10:27

2

Just switching from page.text to page.content should make it work.

Explanation here.

Also see:

edited May 23 '17 at 10:27

Community

1
1

answered Sep 20 '15 at 02:15

alecxe

462,703
120
1,088
1,195

Thank you very much @alecxe – Huy Do Sep 20 '15 at 02:41

How to get unicode string when extract data in Python?

1 Answers1