Vous pouvez utiliser la méthode story.getInfo de l'API Digg. Un de ses arguments possibles est clean_title que vous pouvez analyser à partir du lien dans le flux RSS. Voici un exemple d'implémentation:
import feedparser
import urllib2
from xml.etree import ElementTree
rss_link = 'http://feeds.digg.com/digg/popular.rss'
api_link = 'http://services.digg.com/1.0/endpoint?method=story.getInfo&clean_title=%s'
data = feedparser.parse(rss_link)
for i, e in enumerate(data.entries, 1):
print '%d. Digg link: %s' % (i, e.link)
title = e.link[e.link.rfind('/') + 1 :]
xml = urllib2.urlopen(api_link % title).read()
tree = ElementTree.fromstring(xml)
print '%d. Real link: %s' % (i, tree.find('story').get('link'))
... qui sort:
1. Digg link: http://feeds.digg.com/~r/digg/popular/~3/V58R-d7nd2M/Pakistan_court_bans_Facebook_site
1. Real link: http://news.bbc.co.uk/2/hi/south_asia/8691406.stm
2. Digg link: http://feeds.digg.com/~r/digg/popular/~3/LoF6h1fTtk/Britons_spend_more_webtime_reading_news_than_looking_at_porn
2. Real link: http://www.telegraph.co.uk/technology/news/7740500/Britons-spend-more-web-time-reading-news-than-looking-at-pornography.html
3. Digg link: http://feeds.digg.com/~r/digg/popular/~3/XQUD2tR-qGQ/Sludgy_oil_begins_washing_into_Lousiana_s_coastal_marshes
3. Real link: http://www.washingtonpost.com/wp-dyn/content/article/2010/05/18/AR2010051801676.html?hpid=topnews
4. Digg link: http://feeds.digg.com/~r/digg/popular/~3/4HBB7lvCpoM/Professor_examines_the_complex_evolution_of_human_morality
4. Real link: http://www.physorg.com/news193472479.html
5. Digg link: http://feeds.digg.com/~r/digg/popular/~3/9__2-MVmSp4/How_Are_America_s_Top_Companies_Taxed_Infographic
5. Real link: http://www.mint.com/blog/trends/how-are-americas-top-companies-taxed/
...
thx, je ne savais pas digg a une API. – Timmy