Skip to main content icon/video/no-internet

Web Mining

Web mining is a form of data mining oriented toward analyzing portions of the World Wide Web. This can include information contained within it, as well as information regarding its usage and structure. Data are gathered via automated means, either directly from the web itself or indirectly through server usage logs. It is then analyzed using data mining techniques. The results can be used for a wide variety of purposes, including technical assessments, targeted advertising, market research, military intelligence, and government analysis, among many other possible applications. This entry examines the types and goals of web mining and discusses controversies associated with web mining.

Types and Goals of Web Mining

Web mining can be divided into three general areas: (1) web content mining, (2) web usage mining, (3) and web structure mining (Figure 1). Web content mining analyzes the data hosted on websites. This usually includes text, but it can also include images, audio, video, and other files or data sources. For example, a web mining system could analyze Facebook pages and extract the names, contact information, hobbies, and activities of users. In this way, a large amount of data can be obtained in a relatively short period of time. The process by which these data are harvested is called “scraping,” which may be done using different techniques depending on what is needed for the mining effort. Improvements in analysis and artificial intelligence allow for increasingly thorough examination of a greater variety of data sources.

Web usage mining focuses on observing the usage of web resources. This may include individual users or groups of users. Often, this is accomplished through the use of a tracking cookie, which is a piece of data used for communication between a web browser and a web server. This cookie allows systems to work together in order to monitor and track the sites a user has visited. Often, other information about the visit, such as the time it occurred, the pages visited, and the time between page transitions, is also recorded. The data can then be used to compile information on individuals, as well as aggregate information to help monitor behavior over larger demographics. There are also other approaches, such as analyzing activity logged by the web server itself. Log analysis can provide a more thorough picture of web activity, but generally, the results will apply primarily to a single site or network.

Figure 1 The Relationship Between Types of Web Mining and Data Mining Objectives

None

Source: https://upload.wikimedia.org/wikipedia/commons/6/6c/The_general_relationship_between_the_categories_of_Web_Mining_and_objectives_of_Data_Mining%28English_version%29.png.

Note: Different forms of web mining can yield data on web usage, contents, and the structure of the Internet.

Web structure mining focuses on gathering and analyzing data on the connections (e.g., links) among webpages. This can help understand the technical properties of portions of the web, and often, other information can be inferred from this structure, which is literally shaped like a web in many cases. For example, one can choose a specific website and then assess other websites to see how many times they link to that particular site. A large number of links relative to other websites may suggest greater popularity. Observations such as these are used in a variety of tasks, such as ranking hits for search engine results.

...

  • Loading...
locked icon

Sign in to access this content

Get a 30 day FREE TRIAL

  • Watch videos from a variety of sources bringing classroom topics to life
  • Read modern, diverse business cases
  • Explore hundreds of books and reference titles

Sage Recommends

We found other relevant content for you on other Sage platforms.

Loading