python爬虫爬取淘宝商品信息（selenum+phontomjs）

脚本专栏 2024/10/5 佚名

2 0 1

神剑山庄资源网 Design By www.hcban.com

本文实例为大家分享了python爬虫爬取淘宝商品的具体代码，供大家参考，具体内容如下

1、需求目标 ：

进去淘宝页面，搜索耐克关键词，抓取商品的标题，链接，价格，城市，旺旺号，付款人数，进去第二层，抓取商品的销售量，款号等。

2、结果展示

3、源代码

# encoding: utf-8
import sys
reload(sys)
sys.setdefaultencoding('utf-8')
import time
import pandas as pd
time1=time.time()
from lxml import etree
from selenium import webdriver
#########自动模拟
driver=webdriver.PhantomJS(executable_path='D:/Python27/Scripts/phantomjs.exe')
import re

#################定义列表存储#############
title=[]
price=[]
city=[]
shop_name=[]
num=[]
link=[]
sale=[]
number=[]

#####输入关键词耐克(这里必须用unicode)
keyword="%E8%80%90%E5%85%8B"


for i in range(0,1):

  try:
    print "...............正在抓取第"+str(i)+"页..........................."

    url="https://s.taobao.com/search"+str(i*44)
    driver.get(url)
    time.sleep(5)
    html=driver.page_source

    selector=etree.HTML(html)
    title1=selector.xpath('//div[@class="row row-2 title"]/a')
    for each in title1:
      print each.xpath('string(.)').strip()
      title.append(each.xpath('string(.)').strip())


    price1=selector.xpath('//div[@class="price g_price g_price-highlight"]/strong/text()')
    for each in price1:
      print each
      price.append(each)


    city1=selector.xpath('//div[@class="location"]/text()')
    for each in city1:
      print each
      city.append(each)


    num1=selector.xpath('//div[@class="deal-cnt"]/text()')
    for each in num1:
      print each
      num.append(each)


    shop_name1=selector.xpath('//div[@class="shop"]/a/span[2]/text()')
    for each in shop_name1:
      print each
      shop_name.append(each)


    link1=selector.xpath('//div[@class="row row-2 title"]/a/@href')
    for each in link1:
      kk="https://" + each


      link.append("https://" + each)
      if "https" in each:
        print each

        driver.get(each)
      else:
        print "https://" + each
        driver.get("https://" + each)
      time.sleep(3)
      html2=driver.page_source
      selector2=etree.HTML(html2)

      sale1=selector2.xpath('//*[@id="J_DetailMeta"]/div[1]/div[1]/div/ul/li[1]/div/span[2]/text()')
      for each in sale1:
        print each
        sale.append(each)

      sale2=selector2.xpath('//strong[@id="J_SellCounter"]/text()')
      for each in sale2:
        print each
        sale.append(each)

      if "tmall" in kk:
        number1 = re.findall('<ul id="J_AttrUL">(.*"NULL")

      if "taobao" in kk:
        number2=re.findall('<ul class="attributes-list">(.*"NULL")

      if "click" in kk:
        number.append("NULL")

  except:
    pass


print len(title),len(city),len(price),len(num),len(shop_name),len(link),len(sale),len(number)

# #
# ######数据框
data1=pd.DataFrame({"标题":title,"价格":price,"旺旺":shop_name,"城市":city,"付款人数":num,"链接":link,"销量":sale,"款号":number})
print data1
# 写出excel
writer = pd.ExcelWriter(r'C:\\taobao_spider2.xlsx', engine='xlsxwriter', options={'strings_to_urls': False})
data1.to_excel(writer, index=False)
writer.close()

time2 = time.time()
print u'ok,爬虫结束!'
print u'总共耗时：' + str(time2 - time1) + 's'
####关闭浏览器
driver.close()

以上就是本文的全部内容，希望对大家的学习有所帮助，也希望大家多多支持。

python爬虫,python爬取淘宝商品,python爬取

标签：

python爬虫,python爬取淘宝商品,python爬取

神剑山庄资源网 Design By www.hcban.com

神剑山庄资源网 免责声明：本站文章均来自网站采集或用户投稿，网站不提供任何软件下载或自行开发的软件！如有用户或公司发现本站内容信息存在侵权行为，请邮件告知！ 858582#qq.com

神剑山庄资源网 Design By www.hcban.com

评论“python爬虫爬取淘宝商品信息（selenum+phontomjs）”

暂无python爬虫爬取淘宝商品信息（selenum+phontomjs）的评论...

www.hcban.com 神剑山庄资源网

139,976影音资源

144,792福利资源

21,817软件资源

631,128技术资源

最新文章

何洛洛.2024-别叫醒我（EP）【光羽】【FLAC分

2024/10/5

林忆莲.1996-爱莲说2CD【华纳】【WAV+CUE】

2024/10/5

黄妃.2005-红【亚律】【WAV+CUE】

2024/10/5

刘美麟《同生》[FLAC/分轨][161.95MB]

2024/10/5

群星《前途海量电影原声专辑》[320K/MP3][

2024/10/5

一句话新闻

苹果官宣WWDC 2024！预计会有大批AI功能 - 2024/10/5

3月27日消息，苹果宣布2024年全球开发者大会（WWDC）将于6月10日至6月14日举行，巧合的是，这次大会与端午假期重合。

苹果官方表示：

在线参加 Apple 每年规模最大的开发者盛会。亲眼见证 Apple 最新平台、技术和工具的发布。了解如何创建和改进你的 App 和游戏。与 Apple 设计师和工程师互动交流，与全球开发者社区建立联系。以上活动均免费在线举行。

探索各种新的工具、框架和功能，助力你打造出理想的 App 和游戏。通过视频讲座学习新技能，与 Apple 专家进行一对一会面，以推进你的项目，完善你的构思。

Swift Student Challenge 旨在支持和鼓舞下一代开发者、创作者和企业家。太平洋时间 3 月 28 日，我们将公布今年的获奖者名单。获奖者将有资格参加在 Apple Park 举办的特别活动。我们还会选出 50 名杰出获胜者，他们将受邀前往库比提诺，获得为期三天的非凡体验，包括参加 Apple Park 的特别活动。

python爬虫爬取淘宝商品信息（selenum+phontomjs）

python爬虫,python爬取淘宝商品,python爬取

python中正则表达式的使用方法

python正则表达式爬取猫眼电影top100

评论“python爬虫爬取淘宝商品信息（selenum+phontomjs）”

RTX 5090要首发性能要翻倍！三星展示GDDR7显存

更新日志

友情链接

python爬虫爬取淘宝商品信息（selenum+phontomjs）

python爬虫,python爬取淘宝商品,python爬取

python中正则表达式的使用方法

python正则表达式爬取猫眼电影top100

评论“python爬虫爬取淘宝商品信息（selenum+phontomjs）”

RTX 5090要首发 性能要翻倍！三星展示GDDR7显存

更新日志

友情链接

RTX 5090要首发性能要翻倍！三星展示GDDR7显存