Intelligent focused crawler: Learning which links to crawl

Taylan, Duygu; Poyraz, Mitat; Akyokuş, Selim; Ganiz, Murat Can

Intelligent focused crawler: Learning which links to crawl

Dosyalar

sakyokus_2011.pdf (2.05 MB)

Tarih

2011-06

Yazarlar

Yayıncı

IEEE

Erişim Hakkı

info:eu-repo/semantics/closedAccess

Özet

A web crawler is defined as an automated program that methodically scans through Internet pages and downloads any page that can be reached via links. With the exponential growth of the Web, fetching information about a special-topic is gaining importance. A focused crawler is a web crawler that attempts to download only web pages that are relevant to a predefined topic or set of topics. In order to determine a web page is about a particular topic, focused crawlers use classification techniques. In this study we focus on the classification of links instead of downloaded web pages to determine relevancy. We combine a Naïve Bayes classifier for classification of URLs with a simple URL scoring optimization to improve the system performance. Our results demonstrate that proposed approach performs better.

Açıklama

Akyokuş, Selim (Dogus Author) -- Ganiz, Murat C. (Dogus Author) -- Conference full title: 2011 International Symposium on Innovations in Intelligent Systems and Applications (INISTA 2011) Istanbul, Turkey, 15 - 18 June 2011

Anahtar Kelimeler

Focused Crawler, Link Classification, Machine Learning, Naive Bayes, Turkish Web Pages, URL Optimization

Kaynak

2011 International Symposium on Innovations in Intelligent Systems and Applications (INISTA)

Scopus Q Değeri

N/A

Künye

Taylan, D., Poyraz, M., Akyokuş, S., & Ganiz, M. C. (2011). Intelligent focused crawler: Learning which links to crawl. In 2011 International Symposium on Innovations in Intelligent Systems and Applications (INISTA) (pp. 504-508). Piscataway, NJ: IEEE. https://dx.doi.org/10.1109/INISTA.2011.5946150

Bağlantı

https://dx.doi.org/10.1109/INISTA.2011.5946150
https://hdl.handle.net/11376/2355

Koleksiyon

MF, Bilgisayar Mühendisliği Bölümü, Bildiri & Sunum Koleksiyonu
Scopus İndeksli Yayınlar Koleksiyonu

Detaylı Öğe Kaydı

Intelligent focused crawler: Learning which links to crawl

Dosyalar

Tarih

Yazarlar

Dergi Başlığı

Dergi ISSN

Cilt Başlığı

Yayıncı

Erişim Hakkı

Özet

Açıklama

Anahtar Kelimeler

Kaynak

WoS Q Değeri

Scopus Q Değeri

Cilt

Sayı

Künye

Bağlantı

Koleksiyon

Onay

İnceleme

Ekleyen

Referans Veren