python 之 BeautifulSoup标签查找与信息提取

首页 > 代码库 > python 之 BeautifulSoup标签查找与信息提取

python 之 BeautifulSoup标签查找与信息提取

2024-09-11 23:35:02 216人阅读

一、查找a标签

（1）查找所有a标签

>>> for x in soup.find_all(‘a‘):
    print(x)
    
<a class="sister" href=http://www.mamicode.com/"http://example.com/elsie" id="link1">Elsie</a>
<a class="sister" href=http://www.mamicode.com/"http://example.com/lacie" id="link2">Lacie</a>
<a class="sister" href=http://www.mamicode.com/"http://example.com/tillie" id="link3">Tillie</a>

（2）查找所有a标签，且属性值href中需要保护关键字“”

>>> for x in soup.find_all(‘a‘,href = http://www.mamicode.com/re.compile(‘lacie‘)):
    print(x)

<a class="sister" href=http://www.mamicode.com/"http://example.com/lacie" id="link2">Lacie</a>

（3）查找所有a标签，且字符串内容包含关键字“Elsie”

>>> for x in soup.find_all(‘a‘,string = re.compile(‘Elsie‘)):
    print(x)
    
<a class="sister" href=http://www.mamicode.com/"http://example.com/elsie" id="link1">Elsie</a>

（4）查找body标签的所有子标签，并循环打印输出

>>> for x in soup.find(‘body‘).children:
    if isinstance(x,bs4.element.Tag):        #使用isinstance过滤掉空行内容
        print(x)
        
<p class="title"><b>The Dormouse‘s story</b></p>
<p class="story">Once upon a time there were three little sisters; and their names were
<a class="sister" href=http://www.mamicode.com/"http://example.com/elsie" id="link1">Elsie</a>,
<a class="sister" href=http://www.mamicode.com/"http://example.com/lacie" id="link2">Lacie</a> and
<a class="sister" href=http://www.mamicode.com/"http://example.com/tillie" id="link3">Tillie</a>;
and they lived at the bottom of a well.</p>

二、信息提取（链接提取）

（1）解析信息标签结构，查找所有a标签，并提取每个a标签中href属性的值（即链接），然后存在空列表；

>>> linklist = []
>>> for x in soup.find_all(‘a‘):
    link = x.get(‘href‘)
    if link:
        linklist.append(link)
       
>>> for x in linklist:        #验证：环打印出linklist列表中的链接
    print(x)
  
http://example.com/elsie
http://example.com/lacie
http://example.com/tillie

小结：链接提取 <---> 属性内容提取 <---> x.get(‘href‘)

（2）解析信息标签结构，查找所有a标签，且每个a标签中href中包含关键字“elsie”,然后存入空列表中；

>>> linklst = []
>>> for x in soup.find_all(‘a‘, href = http://www.mamicode.com/re.compile(‘elsie‘)):
    link = x.get(‘href‘)
    if link:
        linklst.append(link)
    
>>> for x in linklst:        #验证：循环打印出linklist列表中的链接
    print(x)
    
http://example.com/elsie

小结：在进行a标签查找时，加入了对属性值href内容的正则匹配内容 <---> href = http://www.mamicode.com/re.compile(‘elsie‘)

（3）解析信息标签结构，查询所有a标签，且每个a标签中字符串内容包含关键字“Elsie”,并将输出结构放入空列表中；

>>> for x in soup.find_all(‘a‘):
    string = x.get_text()
    print(string)
   
Elsie
Lacie
Tillie

python 之 BeautifulSoup标签查找与信息提取

声明：以上内容来自用户投稿及互联网公开渠道收集整理发布，本网站不拥有所有权，未作人工编辑处理，也不承担相关法律责任，若内容有误或涉及侵权可进行投诉：投诉/举报工作人员会在5个工作日内联系你，一经查实，本站将立刻删除涉嫌侵权内容。

联系
我们

首页 > 代码库 > python 之 BeautifulSoup标签查找与信息提取

python 之 BeautifulSoup标签查找与信息提取

看完仍有疑问？有类似问题直接问程序猿