Scrapy: Extract links and text

Question

I am new to scrapy and I am trying to scrape the Ikea website webpage. The basic page with the list of locations as given here.

My items.py file is given below:

import scrapy


class IkeaItem(scrapy.Item):

    name = scrapy.Field()
    link = scrapy.Field()

And the spider is given below:

import  scrapy
from ikea.items import IkeaItem
class IkeaSpider(scrapy.Spider):
    name = 'ikea'

    allowed_domains = ['http://www.ikea.com/']

    start_urls = ['http://www.ikea.com/']

    def parse(self, response):
        for sel in response.xpath('//tr/td/a'):
            item = IkeaItem()
            item['name'] = sel.xpath('a/text()').extract()
            item['link'] = sel.xpath('a/@href').extract()

            yield item

On running the file I am not getting any output. The json file output is something like:

[[{"link": [], "name": []}

The output that I am looking for is the name of the location and the link. I am getting nothing. Where am I going wrong?

@aberna what difference will that make? I'll try that ASAP and no difference. No output. — praxmon, Jan 03 '15 at 09:05
It would follow the scrapy example as in the documentation (http://doc.scrapy.org/en/latest/topics/spiders.html) — aberna, Jan 03 '15 at 09:06

score 23 · Accepted Answer · answered Jan 03 '15 at 09:10

23

There is a simple mistake inside the xpath expressions for the item fields. The loop is already going over the a tags, you don't need to specify a in the inner xpath expressions. In other words, currently you are searching for a tags inside the a tags inside the td inside tr. Which obviously results into nothing.

Replace a/text() with text() and a/@href with @href.

(tested - works for me)

answered Jan 03 '15 at 09:10

alecxe

462,703
120
1,088
1,195

Could you please explain why this works and what I am trying does not? Basically I want to know how and where I was going wrong. Thanks for the answer. It works. :) – praxmon Jan 03 '15 at 09:13
@PrakharMohanSrivastava updated the answer. Sorry, I'm not really good at explaining things :) – alecxe Jan 03 '15 at 09:15

score 5 · Answer 2 · answered Dec 30 '15 at 07:14

5

use this....

    item['name'] = sel.xpath('//a/text()').extract()
    item['link'] = sel.xpath('//a/@href').extract()

answered Dec 30 '15 at 07:14

Ganesh

837
1
11
18

2

use this and try this tend to be poor things to say in an explanation – Drew Dec 30 '15 at 07:20
3

thanks drew, I think these kind explanation goes you up. – Ganesh Dec 30 '15 at 07:52
2

not sure what that means. Trying to help you gain points by good answers. – Drew Dec 30 '15 at 07:57

Scrapy: Extract links and text

2 Answers2

Linked