CROSS-LANGUAGE KNOWLEDGE SHARING MODEL BASED ON ONTOLOGIES AND LOGICAL INFERENCE
WEISEN GUO
Science Integration Program (Human), Department of Frontier Sciences and Science Integration, Division of Project Coordination, The University of Tokyo, 5-1-5 Kashiwa-No-Ha
Kashiwa-Shi, Chiba-Ken 277-8568, Japan
E-mail: gws@scint.dpc.u-tokyo.ac.jp
STEVEN B. KRAINES†
Science Integration Program (Human), Department of Frontier Sciences and Science Integration, Division of Project Coordination, The University of Tokyo, 5-1-5 Kashiwa-No-Ha
Kashiwa-Shi, Chiba-Ken 277-8568, Japan
†E-mail: sk@scint.dpc.u-tokyo.ac.jp
Vast amounts of new knowledge are created on the Internet in many different languages every day. How to share and search this knowledge across different languages efficiently is a critical problem for information science and knowledge management. Conventional cross-language knowledge sharing models are based on natural language processing (NLP) technologies. However, natural language ambiguity, which is a problem even for single language NLP, is exacerbated when dealing with multiple languages. Semantic web technologies can circumvent the problem of natural language ambiguity by enabling human authors to specify meaning in a computer-interpretable form. In particular, description logics ontologies provide a way for authors to describe specific relationships between conceptual entities in a way that computers can process to infer implied meaning. This paper presents a new cross-language knowledge sharing model, SEMCL, which uses semantic web technologies to provide a potential solution to the problem of ambiguity. We first describe the methods used to support searches at the semantic predicate level in our model. Next, we describe how our model realizes a cross-language approach. We present an implementation of the model for the general engineering domain and give a scenario describing how the model implementation handles semantic cross-language knowledge sharing. We conclude with a discussion of related work.
1. Introduction
We live in an age of knowledge explosion. Knowledge sharing can significantly increase social capital (Widen-Wulff et al., 2004). But much of knowledge on the Internet is represented in diverse languages, which limits our ability to share and search knowledge globally. The traditional approach to share knowledge across diverse languages by manually translating each knowledge resource from the original language to all of the other languages is too slow and costly for tasks such as sharing scientific findings between researchers.
Automated cross-language technologies have been developed that use natural language processing (NLP) technologies to extract keywords for matching knowledge resources between different languages. However, NLP-based approaches cannot produce accurate matching results because of the ambiguity of natural language (Hunter and Cohen, 2006). Even thesauri or classification schemata are insufficient (Goldschmidt and Krishnamoorthy, 2008) because they do not support expressions of semantic relationships between keywords or named entities in text. Furthermore, the need to handle multiple languages in cross-language knowledge sharing models exacerbates the problem of natural language ambiguity. Some approaches to decrease the ambiguity have been reported in the literature. For example, Littman et al. (1998) used a latent semantic indexing technique to implement cross-language information retrieval. However, even these sophisticated NLP technologies do not address the fundamental issue of ambiguity in representing knowledge with natural language, an issue that is particularly problematic in a multilingual knowledge sharing situation.
Semantic Web technologies can be used to express knowledge in a computer-interpretable enable matching at a semantic predicate level, e.g. matching of both named entities and predicates stating the semantic relationships between them. Specifically, ontologies constructed in a language such as OWL-DL can represent domain knowledge within a description logic (DL) formalism (www.w3.org/TR/2004/REC-owl-features-20040210). Then DL-based inference can be used in knowledge search to find more useful matching results (Guo and Kraines, 2008).
We present a cross-language knowledge sharing model in this paper that is based on ontologies and logical inference. Using this model, knowledge providers can publish knowledge resources in their native languages, and knowledge seekers can search for knowledge in different languages, thereby enabling cross-language knowledge sharing. Furthermore, both the descriptors of the knowledge resources and the search queries are represented in a form that can be interpreted semantically by a computer, which enables the computer to infer embedded meaning that is implied but not explicitly expressed. Therefore, the knowledge system implementing this model returns matching results represented in diverse languages that should be more accurate than those of conventional keyword based systems because matching is done at the semantic predicate level.
The rest of the paper is organized as follows. In section 2, we review the state of the art of knowledge sharing on the Internet. In section 3, we present the cross-language knowledge sharing model and describe the cross-language method that we have developed to implement the model. We discuss the related work in Section 4 and conclude this paper in Section 5.
2. Knowledge Sharing using Information Retrieval and Semantic Web Technologies
Knowledge sharing is an activity through which knowledge is exchanged among people and/or organizations (Lin, 2007). In this paper, we focus on knowledge existing in explicit digital form on the Internet. A knowledge sharing community consists of two main types of the knowledge users: knowledge providers and knowledge seekers. Community members can be both types: each knowledge user may both provide knowledge resources and seek knowledge resources. The goal of a knowledge sharing system is to return the correct knowledge resources to the knowledge seeker. A global-scale knowledge sharing community will invariably include knowledge users from different countries speaking different languages.
Information Retrieval (IR) technologies can quickly find matching results by matching keywords provided by knowledge seekers with knowledge resources that are represented in natural language. In this approach, the matching system uses automatic techniques such as Natural Language Processing to determine which knowledge resources match the keywords based on the natural language representations of those resources. Conventional cross-language knowledge sharing models are based on these IR technologies. However, the problems of natural language ambiguity and grammatical complexity, which already make it difficult to determine matches with free-text in a single language, become even more serious when dealing with different languages, which results in a rapid decrease in matching precision and recall.
The problem of natural language ambiguity can be addressed by enabling people to create descriptions of knowledge resources in a computer-understandable format. For example, accuracy of matching knowledge resources can be increased by considering predicate-level semantics (Hunter et al., 2008). In particular, Semantic Web technologies, such as ontologies and logical inference, can be used to implement knowledge sharing systems that can match knowledge resources with search descriptions at the level of a grammatical sentence. These systems, such as EKOSS (Kraines et al., 2006) and Annotea (Kahan et al., 2001), are based on a semantic model for matching knowledge resources, which we call Model SEM. In this model, the knowledge providers describe their knowledge resources using computer-interpretable semantic statements instead of natural language. The knowledge seekers also input their queries in a semantic way, rather than just listing keywords. For example, the EKOSS system uses semantic matching methods based on description logics (Kraines et al., 2006; Guo and Kraines, 2008) to match the descriptions of knowledge resources and the search queries. Because the EKOSS implementation of Model SEM supports semantic matching, it can help the knowledge seekers find more correct matching results by reducing the ambiguity in both descriptions of knowledge resources and search queries.
We suggest that the use of Semantic Web technologies to enable people to create computer-interpretable semantic statements describing knowledge resources and requirements could address the issues of ambiguity and grammatical complexity in cross-language knowledge sharing. Based on Model SEM, this paper presents a new model for cross-language knowledge sharing, which we call Model SEMCL. In this model, the knowledge providers describe their knowledge resources by creating computer-interpretable semantic statements using their preferred language. In the same way, the knowledge seekers use their preferred language to describe their queries. The system is able to infer semantic matches between the descriptions and the queries, and it then displays the matching results in the preferred language of the knowledge seeker (Fig. 1).
The following section gives the details of our proposed Model SEMCL.
3. The Cross-Language Knowledge Sharing Model SEMCL
Fig. 2 shows the framework of Model SEMCL supporting three hypothetical languages: La, Lb, and Lc. Model SEMCL has four levels: the core semantic matching level, the language level, the user interface level, and the human level. At the center of Model SEMCL is the domain ontology, which is comprised of classes and properties together with labels in plain text. Translations of the class and property labels in the domain ontology to the other languages (Ontology(La), Ontology(Lb), and Ontology(Lc)) are made by a human or machine translator before running the knowledge sharing system.
At the human level, there are two types of users: Knowledge Providers and Knowledge Seekers. Knowledge Providers use the domain ontology to create computer-interpretable semantic descriptions of their knowledge resources, descriptions that are populated by instances of the ontology classes together with properties that describe specific relationships between those instances. Because these descriptions are free from the ambiguity of natural language and grounded in the logic supported by the ontology, they can be used by computers for inference (Guo and Kraines, 2008). The language level of Model SEMCL supports the cross-language sharing and searching. The user interface level of Model SEMCL provides multi-language graphic user interfaces (GUIs), which users can use to provide or seek knowledge in their preferred language.
In Fig. 2, a Knowledge Provider, who prefers to use language La, uses the GUI in language La (GUI(La)) to create a semantic description (Statement + La) of her/his knowledge resource in language La. A Knowledge Seeker, who prefers to use language Lc, uses the GUI in language Lc (GUI(Lc)) to create a semantic query (Query + Lc) in language Lc. The matching results produced by the Inference Engine at the core semantic matching level (Result) are augmented with the language information for Lc to create (Result + Lc), which is shown to the Knowledge Seeker in GUI(Lc).
To make this kind of cross-language searching possible, each knowledge description has two parts: semantic statement (Statement) and language information. Each search query also has two parts: semantic query (Query) and language information. The language information is maintained in the language level. The semantic statement and semantic query go into the core level to be matched by the Inference Engine, which uses reasoning in the supported logic as well as optional rule-based reasoning to match all the available semantic statements with each semantic query. When the Inference Engine finds some matching results (Result), it returns them to the Knowledge Seeker. Language information is added to the matching results when it goes through the language level to the user interface of the Knowledge Seeker. In summary, Model SEMCL uses ontologies to handle the ambiguity of natural language, logical inference for semantic matching, and the method of separating language from semantics to handle the cross-language issue. The following subsections give the details for each of these techniques.
3.1. Knowledge representation and search
In Model SEMCL, there are two kinds of knowledge. The first kind is the domain knowledge: the basic concepts and their relationships in the targeted knowledge domain. The second kind is the knowledge that the Knowledge Providers want to share, which is described using the first kind of knowledge. In Model SEMCL, the first kind of knowledge must be created prior to the operation of the knowledge sharing system and kept relatively stable. It should also have sufficient detail to represent the second kind of knowledge, which makes up the contents of the knowledge base in Model SEMCL.
We have created an implementation of Model SEMCL for the domain of engineering knowledge. In our implementation, the first kind of knowledge is represented by using an OWL-DL ontology that we have created for that domain. There are five main classes in the ontology – substances, activities, physical objects, events, and classes of activities (actors and spatial locations are special kinds of physical objects) – as well as several properties that can be used to specify relationships between the classes or instances of the classes (Fig. 3). For example, an instance of the class “activity” can have a relationship with an instance of the class “class of activity” using the property “has activity class”. Each main class is divided into subclasses to represent more specific concepts from the engineering domain.
The second kind of knowledge, knowledge shared by the Knowledge Providers, is represented using the classes and properties provided in the first kind of knowledge. Specifically, the entities described by each piece of shared knowledge are represented as instances of ontology classes, and the specific relationships that are described between those entities are represented using ontology properties. For example, consider the following accident report:
“The central region of Seongsu Bridge, which was built in the capital city Seoul city in Korea, suddenly collapsed on October 21, 1994. A diesel bus fell, and several people were killed. According to the investigation after the accident, the collapse was caused by fractures in the steel girders of the bridge.”
The knowledge that is expressed in this report can be represented using the domain ontology as shown in Fig. 4.
In Model SEMCL, all knowledge resources are represented as knowledge descriptions in this way. When the Knowledge Seekers want to find knowledge resources, they create semantic queries, also based on the domain ontology, and send them to the Inference Engine (an example is given in section 3.3).
Upon receiving a semantic query, the Inference Engine matches it with all the available semantic statements using a DL reasoner. First, the Inference Engine loads the domain ontology to the knowledge base. Then it loads the semantic statement for one knowledge resource. Finally, it evaluates the semantic query against the knowledge base that now contains the ontology and the statement. If each ontology class in the query can be mapped to an instance in the knowledge base subject to the properties specified for that class in the query, then the semantic statement that was loaded to the knowledge base is said to match with the query (an example is given in section 3.3).
3.2. Cross-language Knowledge Sharing
Model SEMCL handles the cross-language issue by separating the language information from the semantic statement in the knowledge description that is created by the Knowledge Provider. Only the semantic statement is used to obtain the matching result. The language information is added to the semantic statement of the matching results before showing them to the Knowledge Seeker. Because Model SEMCL uses a domain ontology instead of natural language to represent the knowledge, it is easy to separate the language information from the semantic statement.
Fig. 5 shows the overall Model SEMCL cross-language mechanism for two languages (La and Lb). In a real application, the number of languages can be more. The domain ontology and the user interface are created to form the infrastructure level and translated into each of the supported languages before running the knowledge sharing system. The interface language and ontology language are paired. In other words, if the interface language is changed to La, then the ontology language is changed to La automatically.
The Knowledge Provider uses the interface and ontology in her/his preferred language to create her/his knowledge descriptions. For example, if the preferred language is La, then “GUI(La)” and “Ontology(La)” are used, and the knowledge description “Statement + La” is created. The “Statement” is just the semantic information that remains when the language information in language La is removed. In the same way, the Knowledge Seeker uses the interface and ontology in her/his preferred language to create her/his search query. For example, if the preferred language is Lb, then “GUI(Lb)” and “Ontology(Lb)” are used, and “Query + Lb” is created. The “Query” is the part that remains when the language information in language Lb is removed. The Inference Engine evaluates matches between the Statements and Queries using the ontology. If “Query” matches with “Statement”, then the matching result “Result” is created by the Inference Engine. Because the preferred language of the Knowledge Seeker is Lb, the language information for Lb is added to create “Result + Lb”, which is displayed to the Knowledge Seeker.
3.3. Scenario
Here, we illustrate how Model SEMCL works by using a scenario involving three Knowledge Providers – Jane, Hideo, and Zhang – who are sharing knowledge on a knowledge sharing system that supports three languages: English, Japanese and Chinese.
In our scenario, Jane is the person who provided the knowledge for the article about the accident described in section 3.1, and she prefers using English. She accesses the Model SEMCL knowledge sharing system and selects the English user interface to create a description for this knowledge resource (see Fig. 4). She adds the URL of the original news report to the description so that anyone finding this description to be of interest can access the knowledge resource (the news report) for details.
Hideo is a scientist studying failure knowledge who prefers using Japanese. He created a video to explain the failure mechanism behind an airplane accident in Israel that he wants to share. He selects the Japanese user interface to create a description for this video, linked to the URL of the video. The corresponding English description is:
“On October 4, 1992, soon after the take-off of a Boeing 747 cargo air transport of El-Al Israel Airlines, the two engines on the right wing dropped off, and the air transport went out of control, finally colliding into the apartment building. Thirty-nine apartment habitants were killed in this accident. The cause of engine separation during take-off phase was fatigue failure of the pylon fuse pin.”
Zhang is a vehicle engineer and a fan of car racing who prefers using Chinese. After learning of the failure of the China A1 Racing Team vehicle in the Indonesian A1 Grand Prix finals station, he created some illustrations to show the reason of transmission failure from his professional perspective that he wants to share. He selects the Chinese user interface to create a description for these illustrations and links the description to the URL of the illustrations. His description of the content of the illustrations in Chinese is as follows:
“中国A1赛车队于2006年2月12日参加了A1大奖赛印尼站决赛。其间,江腾一驾驶的赛车失去控制停在了赛道上。事故原因被认为是经过长期超强度运转的引擎老化引起的。”
The corresponding English translation is:
“The China A1 racing team participated in the Indonesian A1 Grand Prix finals station on February 12, 2006. The vehicle of Jiang Tengyi stopped on the track. The cause of the malfunction was considered to be the aging of the engine after a long time of highly intense use.”
The description created by Zhang is shown in Fig. 6.
A Knowledge Seeker named Helen, who prefers using English, is composing her thesis about mechanical failures of transportation machines. In order to find some more knowledge, she wants to utilize the knowledge sharing system. She selects the English user interface to create her query for “a disaster caused by the mechanical failure of a machine artifact that is part of a transportation device” (see Fig. 7).
Helen sends her query to the Inference Engine of the Model SEMCL knowledge sharing system and waits for search results. The Inference Engine compares the semantic part of the query with the semantic statement part of Jane’s description, Hideo’s description and Zhang’s description. Hideo’s description and Zhang’s description match with Helen’s query (Fig. 8). However, Jane’s description does not match. Even though her description does mention a disaster event, an activity, a transportation device, a machine artifact, and a mechanical failure, the object with the mechanical failure is a part of the road bridge, not a transportation device. This example demonstrates how Model SEMCL supports cross-language knowledge sharing based on matching at the semantic predicate level.
4. Related Work
In section 2, we considered some other models for knowledge sharing. In this section, we compare Model SEMCL with some closely related work from the literature.
Littman et al., (1998) used Latent Semantic Indexing (LSI) to retrieve cross-language documents automatically. They treated a set of dual-language documents as training documents to create a dual-language semantic space in which terms from both languages are represented. Standard mono-lingual documents are represented as language-independent numerical vectors in this semantic space, so queries in either language can retrieve documents in either language without the need to translate the query. The LSI method is based on keywords. The semantic space contains the dual-language terms that form an index for speeding up the retrieval. However, the semantic space does not support classifications and relationships, so the LSI method cannot support retrieval based on matching at the semantic predicate level.
Diaz-Galiano et al., (2008) used the Medical Subject Headings (MeSH) to expand queries in the task of multilingual image retrieval. The expansion consists of searching for terms from the topic query in the MeSH vocabulary and adding similar terms. MeSH has a hierarchical structure that provides a consistent way to retrieve information using different terms for the same concepts. However, the MeSH structure does not contain typed relationships. So, the MeSH-based method also does not support retrieval at the semantic predicate level.
Wang et al., (2004) used a Publish/Subscribe system to share the knowledge. In their system, the knowledge descriptions, which they call Events, are represented with RDF. They also used a domain ontology as the domain basic knowledge which specifies the concepts involved in the Events, the relations between them, and the constraints on them. Their knowledge sharing model can be considered as an example of Model SEM.
Kraines et al., (2006) used semantic web technologies to share expert knowledge. The basic knowledge of the domain is presented to knowledge users as domain ontologies. The Knowledge Providers create knowledge descriptions for their knowledge resources, and the Knowledge Seekers create queries to search for knowledge of interest to them. Therefore, this system is also an implementation of Model SEM.
In summary, while the first two related research works support cross-language knowledge sharing, because they do not handle semantics directly, the accuracy of the matching results is limited. The last two related research works support the semantic predicate level search but do not handle the cross-language issue. Model SEMCL is a new contribution that uses the Model SEM approach to support cross-language knowledge sharing at the semantic predicate level.
5. Discussions and Conclusion
In today’s age of information explosion, vast amounts of new knowledge are generated every day in a diversity of languages. How to share and search this knowledge efficiently is one of the most important problems in the information science community. Conventional cross-language knowledge sharing models that are based on natural language processing technologies suffer from the exacerbated effect of ambiguity and grammatical complexity over multiple languages. This paper began with an analysis of the task of knowledge sharing in the Internet environment. An approach to matching knowledge resources using Semantic Web Technologies, called Model SEM, was identified. A new model based on Model SEM, SEMCL, was then proposed for cross-language semantic sharing and searching of knowledge resources. We introduced the framework of our proposal for Model SEMCL, focusing on the knowledge representation and search aspects. We then used a scenario based on an implementation of Model SEMCL for general engineering knowledge to demonstrate how Model SEMCL supports disambiguation, semantic predicate level matching, and cross-language sharing. The original contribution of this work is the creation of a cross-language knowledge sharing model that uses Semantic Web technologies to enable searches across multiple languages at a semantic predicate level.
Our implementation of Model SEMCL is accessible on the EKOSS website (www.ekoss.org). Currently, the EKOSS knowledge sharing system supports three languages: English, Japanese and Chinese. To date, a number of different users of EKOSS have used their preferred languages to create semantic statements to describe their knowledge resources.
6. Acknowledgments
The authors thank the President’s Office of the University of Tokyo for funding support.
References
Diaz-Galiano, M.C., Garcia-Cumbreras, M.A., Martin-Valdivia, M.T., Montejo-Raez, A., and Urena-Lopez, A. (2008). “Integrating MeSH Ontology to Improve Medical Information Retrieval”, In Peters, C., et al. (Eds.): CLEF 2007, LNCS 5152, 601-606.
Goldschmidt, D.E., and Krishnamoorthy, M. (2008). “Comparing keyword search to semantic search: a case study in solving crossword puzzles using the GoogleTM API”, Software-Practice & Experience, 38(4), 417-445.
Guo, W., and Kraines, S. (2008). “Explicit Scientific Knowledge Comparison Based on Semantic Description Matching”, American Society for Information Science and Technology 2008 Annual Meeting, Columbus, Ohio.
Hunter, L., and Cohen, K.B. (2006). “Biomedical Language Processing: What’s Beyond PubMed?”, Molecular Cell, 21, 589-594.
Hunter, L., Lu, Z., Firby, J., Baumgartner Jr, W.A., Johnson, H.L., Ogren, P.V., and Cohen, K.B. (2008). “OpenDMAP: An open source, ontology-driven concept analysis engine, with applications to capturing knowledge regarding protein transport, protein interactions and cell-type-specific gene expression”, BMC Bioinformatics, 9:78, doi:10.1186/1471-2105-9-78.
Kahan, J., Koivunen, M.R., Prud’Hommeaux, E., and Swick, R.R. (2001). “Annotea: An Open RDF Infrastructure for Shared Web Annotations”, Proceedings of the WWW10 International Conference, Hong Kong, 623-632.
Kraines, S., Guo, W., Kemper, B., and Nakamura, Y. (2006). “EKOSS: A Knowledge-User Centered Approach to Knowledge Sharing, Discovery, and Integration on the Semantic Web”, ISWC 2006, 5th International Semantic Web Conference, LNCS 4273, 833-846.
Lin, H.F. (2007). “Effects of extrinsic and intrinsic motivation on employee knowledge sharing intentions”, Journal of Information Science, 33(2) 2007, 135-149.
Littman, M.L., Dumais, S.T., and Landauer, T.K. (1998). “Automatic cross-language information retrieval using latent semantic indexing”, In Grefenstette, G., editor, Cross-Language Information Retrieval, chapter 5. Kluwer Academic Publishers, Boston.
Wang, J., Jin, B., and Li, J. (2004). “An Ontology-based Publish/Subscribe System”, In Jacobsen, H.A., (Ed.): Middleware 2004, LNCS 3231, 232-253.
Widen-Wulff, G., and Ginman, M. (2004). “Explaining knowledge sharing in organizations through the dimensions of social capital”, Journal of Information Science, 30(5) 2004, 448-458.
顯示具有 Semantic Technologies and Web 2.0 Tools 標籤的文章。 顯示所有文章
顯示具有 Semantic Technologies and Web 2.0 Tools 標籤的文章。 顯示所有文章
2009年12月9日 星期三
SOCIAL BOOKMARKING AND TAGGING BEHAVIOR
SOCIAL BOOKMARKING AND TAGGING BEHAVIOR: AN EMPIRICAL ANALYSIS ON DELICIOUS AND CONNOTEA
HELEN S. DU
Department of Computing, The Hong Kong Polytechnic University,
Hung Hom, Kowloon, Hong Kong
E-mail: cshelen@inet.polyu.edu.hk
SAMUEL K.W. CHU1 and FLORENCE T.Y. LAM2
Division of Information Technologies, Faulty of Education, The University of Hong Kong1
Department of Statistics and Actuarial Science, The University of Hong Kong2
E-mail: samchu@hkucc.hku.hk1 and florentia.lam@gmail.com2
Social bookmarking services have shown themselves as common and popular Internet tools by successfully acquiring millions of users, with Delicious being one of the most popular social bookmarking services to the public. While Delicious is used mainly for general purposes, Connotea, another social bookmarking site that primarily serves academic and scientific interests, has become equally popular among researcher groups. This paper attempts to analyze and compare users’ bookmarking and tagging behavior in Connotea and Delicious. The results show that there is a distinctive difference in usage behavior among these two groups of users. Delicious users create bookmarks more frequently than Connotea users, but Connotea users tend to use more distinctive tags for their bookmarks than Delicious users. Moreover, our result from the analysis indicates that the number of bookmarks created is a significant predictor of the quantity of tags used. This study is a starting point from which to explore the reasons behind the difference in social bookmarking and tagging behavior among different usage orientation groups.
1. Introduction
Social bookmarking sites have successfully acquired millions of online users in the recent years. A social bookmarking site provides the service for storing, sharing and discovering a collection of web bookmarks with the help of user-generated taxonomies (folksonomies). The term folksonomy came from the words taxonomy and folk, which is used to name the growing phenomenon of users collaboratively creating and managing metadata by tagging pieces of digital information with their own searchable keywords (Dye, 2006). Folksonomy is also known as social tagging. For simplicity, this paper uses the word tagging to refer to social tagging. In recent years, research has been done on tagging and folksonomy to understand folksonomic patterns (Al-Khalifa, & Davis, 2007) and their trends (Hotho, Jäschke, Schmitz, & Stumme 2006). Very few articles have been written on social bookmarking. The ones that have been published mainly give definitions on social bookmarking and its related concepts (Golder & Huberman, 2006) and provide general discussions on social bookmarking tools (Hammond, Hannay, Lund and Scott, 2005; Lund, Hammond, Flack and Hannay, 2005).
Delicious, created by Joshua Schachter in 2003, is one of the most popular social bookmarking service sites. Until now, it has acquired over 5 million users and 150 million bookmarked URLs [http://en.wikipedia.org/wiki/Delicious, viewed June 8, 2009]. Diigo is another similar social bookmarking service like Delicious, which allows users to bookmark and tag web pages. It also contains the functions to highlight and paste sticky notes, so that users can create their personal notes on the content of WebPages they visit. While Delicious and Diigo are social bookmarking tools that serve general purposes, some other social bookmarking tools such as CiteULike and Connotea are being targeted at academic and scientific purposes (Gordon-Murnane, 2006). These two social bookmarking service sites operate in the similar style as of Delicious, in addition to capturing the bibliographic information from scientific articles and journals.
From a preliminary analysis using a data mining technique (association rule mining), we find a notable difference in the usage pattern of the top 100 frequently used tags between Connotea and Delicious. As shown in Figure 1, Delicious has more diversified relationships of tags than Connotea.
Fig. 1. Association rules found in Connotea and Delicious with .1% minimum support and 50% confidence.
Note. Each element in the graph represents a tag. The arrow connecting any tags implies their association relationship. For example, one user uses the tag “protein” may also use “human” to bookmark the same URL (“unidirectional relationship”); one user uses the tag “protein” will also use “RNA” and vice versa (“bidirectional relationship”).
Based on the above preliminary analysis, we suspect that there exists different tagging behavior among the users of Delicious and Connotea that requires further investigation. This investigation is worthwhile as Delicious and Connotea are two very popular social bookmarking sites that serve different target user groups (general-purpose users and researchers respectively). So far, no published articles on social bookmarking have yet attempted to examine and compare social bookmarking services with different orientations. Therefore, this study investigates both Delicious and Connotea, and examines their similarities and differences in users’ bookmarking and tagging behavior.
2. Literature review
2.1 What is social bookmarking?
When people start to rely on the use of Internet, one of the greatest challenges they face is to remember and retrieve items that they have previously found to be useful. The common approach to arrange information on the Web is through the use of personal bookmarks (Millen, Feinberg, & Kerr, 2005). The desire to share information among communities has led to the development of shared bookmarking systems. Social bookmarking is a way to locate, classify and share Internet resources through the use of shared lists of user-created Internet bookmarks. Social bookmarking tools allow users to create tags for bookmarks they saved, and organize all users’ tags so that users can search and browse the tags to find out not only their own bookmarks but also other users’ bookmarks.
2.2 Social bookmarking services
Delicious [http://delicious.com], being the most commonly and widely used social bookmarking service among online users, is a server-based system with a simple-to-use interface that allows users to organize and share bookmarks on the Internet. Hotho et al. (2006) saw such service yields benefits for each individual user (e.g. organizing one’s bookmarks in a browser-independent, persistent fashion) without too much overhead, while there is a proliferation of resources in the Web that makes it difficult to remain up-to-date and to keep track on documents that are related to one’s own area of interest. The usage of social bookmarking services indicates that folksonomy-based approaches seem to be the solution to overcome this difficulty.
A review on social bookmarking services, particularly Connotea [www.connotea.org], and their advantages were summarized by Hammond et al. (2005) and in its companion paper by Lund et al. (2005). Connotea makes sharing among personal collections of resources much easier than before. Instead of placing materials hierarchically in folders, Connotea allows users to create simple tags to the bookmarks. Tagging allows the organization of bookmarks to be more flexible, multi-faceted and spacious. Furthermore, all bookmarks posted by these tools are visible to registered users, which take the concept of sharing to a higher level and benefit “not just from the ease with which it allows explicit sharing with friends and colleagues, but from many users storing their bookmarks in the same space” (Lund et al. 2005, p. 4).
2.3 Advantages/ Disadvantages of social bookmarking
Several researchers (Hotho et al, 2006; Laura, 2006; Menchen, 2005; Millen, 2006) have identified various advantages of social bookmarking. Social bookmarking services provide online storage that facilitates a single repository (Menchen, 2005). They allow an individual to create personal collections of bookmarks and facilitate sharing of bookmarks (Menchen, 2005; Millen, 2006). Social bookmarking services also allow users to create tags, which help organize and categorize users’ collection of information (Gordon-Murnane, 2006; Millen, 2006). By searching or browsing specific tags, users can retrieve all the items bookmarked by other users, which are tagged with specific keywords (Gordon-Murnane, 2006). Besides, bookmarks become portable since they are web-based. Users can access to the links and sites from computers, in contrast to the traditional way of accessing bookmarks from dedicated computers (Gordon-Murnane, 2006).
However, tags lack hierarchy so that a search of a specific term will only yield results on that term and not provide the full body of related terms that might be relevant to the user’s information needs and goals (Gordon-Murnane, 2006). Tags can be considered uncontrolled vocabulary shared across the entire social bookmarking system, and they have inherent ambiguity as different users apply terms to documents in different ways. Tags have no synonym control in the system, words with similar intended meanings, plural and singular forms will also appear in the system (Mathes, 2004). Al-Khalifa and Davis (2007) conducted a study to analyze tags in Delicious and classify them into three groups. They found that the non-standardized forms of tags make it difficult to capture them in a general form. For example, there are spelling variations, compound tags with different combination etc.
2.4 Users’ tagging behavior
Hotho et al. (2006) showed how topic-specific trends can be discovered in Delicious, one of the folksonomy-based systems. They collected users’ profiles and tags in Delicious within a time frame, and analyzed the tags to discover topic-specific trends. They argued that their analysis can be done regardless of the types of the underlying resources which would make folksonomies interesting for multimedia applications.
Kipp’s (2007) study found that a surprising number of tags used in the three social bookmarking tools examined (Delicious, Connotea and Citeulike) were not subject-related. These non subject-related tags can be classified into two groups: affective tags and time and tasks related tags. These behaviors have suggested that “users appear to want to store more than just the subject of the documents they are bookmarking” (Kipp, 2007, p. 3). They revealed specifically the users’ desire to “express an emotional connection to the document” and to “attach personal information management information to documents” (Kipp, 2008, p. 5).
Golder and Huberman (2005, p. 198) investigated the usage patterns of the tagging system in order to “identify regularities in user activity, tag frequencies, kinds of tags used, bursts of popularity in bookmarking and the stability in the relative proportions of tags within a given URL”. It is found that the frequencies of tagging vary much among users. In particular, there is no strong connection between the length of the user account’s existence and the number of days the user takes to create one or more bookmarks. Similarly, the number of bookmarks created by users has very little association with the number of tags used in each bookmark as well. However, a user’s tagging behavior could possibly be used to reflect his or her development of interests. For instance, as a tag grows steadily over time, it might indicate the user’s continual interest in that particular subject. On the other hand, if one tag suddenly grows rapidly, it might reveal the user’s newfound interest (Golder & Huberman, 2005, p. 202).
Other trends in social bookmarking are also observed, such as the time taken for an URL to reach its peak popularity. Although it is found that a majority of the URLs would indeed reach their peak of popularity as soon as they are introduced in Delicious, others have actually taken longer time before they are “rediscovered” and experience a sudden jump (Golder & Huberman, 2005, p. 204-205). In addition, it is often believed that social bookmarking is chaotic, unstructured, and imprecise because the collection of tags depends on each individual’s personal preference and level of knowledge. However, Golder and Huberman’s study claims that each tag’s frequency for a particular URL is a nearly fixed proportion of the total frequency of all tags used. More interestingly, this stability becomes apparent after fewer than 100 bookmarks (Golder & Huberman, 2005, p. 205-206).
2.5 Research Gap
Golder and Huberman (2006) analyzed the structure of collaborative tagging systems of Delicious and discovered regular patterns in user’s activity and tag frequencies. He expected that such findings could be applied to other similar tagging systems. However, little has been done regarding his claim. Therefore, it is worthwhile to conduct a research using similar methodology to that of Golder and Huberman (2006), but with a social bookmarking service designed for scientific purposes, such as Connotea.
The preliminary investigation using data mining also suggests that there is different usage pattern between the two sites – users of Delicious create tags for all kinds of topics while Connotea users create tags more for research purposes. Therefore, it is meaningful to study the users’ tagging behavior in Delicious alongside Connotea to find out if there are any significant similarities and differences between social bookmarking services of different orientations.
3. The Study
3.1 Data Collection and Sampling
Analysis is conducted on two sets of data (one for Delicious and one for Connotea) which contain the activity profiles of some of its users. A program is written using Java 1.4.2 to retrieve data that query both databases. The ways to decide which user should be included are slightly different between Delicious and Connotea because Connotea has a smaller user base than Delicious. For Delicious, the user’s history of activity profile is collected and analyzed if one or more of his/her bookmarks reached the top 500 most popular bookmarks on Feb 16, 2009 (Midnight GMT). This information is found through the public RSS feeds of the ‘popular’ page and then a portion of the website from the user’s activity profile is crawled. For Connotea, the list of users is compiled by examining at the bookmarks posted between April 1, 2008 (Midnight GMT) and May 31, 2008 (23:59 GMT).
This time period is chosen because of two reasons: 1) Most of the researchers and lecturers would have their summer break in June, it would be best to set the time frame before summer to represent the use of Connotea among researchers; 2) the user base for Connotea is smaller than Delicious and a longer period is needed to have a similar amount of bookmarks tagged. Based on the list of users who bookmarked in the aforementioned period, all the bookmarks they have posted since the activation of their accounts till Feb 17, 2009 (0:00:00 GMT) were retrieved. This procedure stopped when over 200,000 bookmarks from both databases were collected. As a result, we collected 5454 users’ data from Connotea and 440 users’ data from Delicious. Note here that the sample size of Connotea is over twelve times more than that of Delicious because Delicious users on average create much more bookmarks than Connotea users (see Table 1).
Table 1. Descriptive Statistics of user profile
Delicious Connotea
Mean 1404.175 49.724
Median 711.500 4
Min 3 1
Max 19100 15067
Note. The values are presented in number of bookmarks
3.2 Data Analysis
This analysis studies the bookmarking and tagging behavior of Connotea and Delicious respectively. The result between Connotea and Delicious is also compared to examine any significant differences.
3.2.1 The descriptive statistics of users’ bookmarking behavior
From the statistics (see Table 2), it is found that the bookmarking behaviors of Connotea and Delicious users differ significantly. It is shown from the standard deviation value and range value that users of Delicious do not deviate a lot in their bookmarking behavior. All users in the sample pool create at least one bookmark every month. However, some Connotea users may only create one bookmark every two years.
Table 2. Descriptive Statistics of users’ bookmarking behavior of Connotea and Delicious
Delicious Connotea
Mean 1.485 8.394
Standard Deviation 2.356 23.614
Minimum 0 0
Maximum 24.221 676.863
Note. The values are present in days.
3.2.2 Relationship between the length of the account since activation and the total number of bookmarks they created
In this regression analysis, both Connotea and Delicious’ data are best fitted into the exponential model. See Table 3 for the exponential regression results.
Delicious. As a result, the length of the user account since activation is a significant predictor of users’ bookmarking behavior variance (F¬ = 74.323, p < .05). In other words, the length of the user account since activation accounts for 14.3% of users’ bookmarking behavior variance. Therefore, it was likely for Delicious users who have created an account for a longer time to create more bookmarks (β = .381, p < .05).
Connotea. The length of the user account’s existence since activation is a significant predictor of users’ bookmarking behavior variance (F¬ = 5806.223, p < .05). In other words, the length of the user account’s since activation accounts for 51.6% of users’ bookmarking behavior variance. Connotea users who have created an account for a longer time were therefore more likely to create more bookmarks (β = .718, p < .05).
Table 3. Univariate Regressions: Delicious versus Connotea
(Dependent Variable: number of bookmarks created)
R2 df F Sig. β t
Delicious a .145 1 74.323 .000 .381 8.621
Connotea a .516 1 5806.223 .000 .718 76.199
a predictor: the length of the user account’s existence in days
In addition, a two-way between-subjects analysis of variance is conducted to compare the effect of the length of account since activiation on number of bookmarks created in Delicious and Connotea. The between-subjects factors are length of the existence of the account and nominal value with two levels (1 or 0 representing Connotea or Delicious, respectively). A significant interaction effect was found, F (1, 5890) = 1071.98, p < .000. The positive coefficient of interaction term (0.0039) suggests the length of account’s existence has a greater effect on number of bookmarks created in Connotea than Delicious.
3.3.3 Relationship between the number of bookmarks a user creates and the number of tags they use in those bookmarks.
In this analysis, the overall relationship between the number of bookmarks users created and the number of tags they used in those bookmarks will be analyzed. The lower end of the scale (users keeping fewer than 30 bookmarks) and the upper end (users keeping more than 500 bookmarks) will also be examined. Both Connotea and Delicious’ data are best fitted into the exponential model for regression analysis. Table 4 gives the exponential regression results.
Delicious. As a result, the number of bookmarks a Delicious user creates is a significant predictor of users’ tagging behavior variance (F¬ = 1207.6, p < .05). In other words, the number of bookmarks a user creates is able to account for 73.7% of users’ tagging behavior variance. Users who create more bookmarks are also likely to use more tags for those bookmarks (β = .858, p < .05). The relationship is found to be weaker at the lower end of the scale with users having fewer than 30 bookmarks (R = .652, R2 = .425, p < .05), but stronger at the upper end with users having more than 500 bookmarks (R = .716, R2 = .512, p < .05).
Connotea. The number of bookmarks a Connotea user creates is a significant predictor of users’ tagging behavior variance (F¬ = 38941.46, p < .05). The results suggest that the number of bookmarks a user creates is able to account for 87.7% of users’ tagging behavior variance. Users who create more bookmarks will also likely to use more tags for those bookmarks (β = .937, p < .05). It is stronger at the lower end of the scale, with users having fewer than 30 bookmarks (R = .845, R2 = .715, p < .05) but comparatively weaker at the upper end, with users having more than 500 bookmarks (R = .825, R2 = .681, p < .05).
Table 4. Univariate Regressions: Delicious Versus Connotea
(Dependent Variable: number of tags they use in bookmarks)
N R2 df F β t
Delicious a
Overall
442
.737
1
1207.6**
.858
34.751
Lower end (< 30 bookmarks) 14 .425 1 9.604** .652 3.099
Upper end (>500 bookmarks) 277 .512 1 289.76** .716 17.022
Connotea a
Overall
5455
.877
1
38941.466**
.937
197.336
Lower end (< 30 bookmarks)
Upper end (>500 bookmarks) 4334 .715 1 10848.069** .845 104.154
92 .681 1 194.527** .825 13.947
a predictor: number of bookmarks a user creates
**p < .01.
In addition, a two-way between-subjects analysis of variance is conducted to compare the effect of the number of bookmarks created on number of tags used in Delicious and Connotea. The between-subjects factors are number of bookmarks created and nominal value with two levels (1 and 0 represents Connotea and Delicious respectively). A significant interaction effect was found, F (1, 5884) = 5.79, p < .05. The negative coefficient of interaction term (-0.054) suggests the number of bookmarks created has a greater effect on number of tags used in Delicious than Connotea.
3.3.3 Proportion of unique tags over all tags used per user account
Figure 3 shows a significant difference in tagging behavior between Connotea and Delicious users. Almost half of Connotea users use mainly distinctive tags to organize or classify their bookmarks, while most of Delicious users use comparatively less unique tags for organization and classification.
Fig. 3. Proportion of unique tags over all tags used per user account in Connotea and Delicious
Note. 6 users from Delicious are excluded since they do not have any tags.
4. Discussion and Implications
Golder and Huberman (2006) expected that their findings on Delicious can be applied to other similar tagging systems. However, our findings suggest that Delicious and Connotea are quite different from one another. Sections 3.2.1 and 3.2.2 analyze the bookmarking behavior of Delicious and Connotea users, and show a distinctive difference between them. The reason behind this finding can be explained by the different user orientations of these two social bookmarking services. Since Delicious is used for general purposes, its users may use Delicious to bookmark websites related to personal interests or entertainment etc., which may occur frequently. Connotea, on the other hand, is a social bookmarking service that caters specifically to the management of scientific references. Users of Connotea mainly use its tools for doing research or for other academic purposes. Once they complete their assignments or research projects, they usually stop using Connotea until they start another project.
Sections 3.2.3 and 3.2.4 examine the tagging behavior of Delicious and Connotea users. Section 3.2.3 shows that both groups of users will use more tags if they have more bookmarks. It is noted that Connotea requires users to use at least one tag per bookmark while Delicious has no such restriction. These findings may suggest that tagging is a useful tool for users to manage, classify and organize bookmarks so that users are willing to create tags on bookmarks even if they are not required to do so. However, users of Delicious and Connotea have different tagging behavior on the bookmarks they created. Section 3.2.4 shows that Connotea users are more likely to use distinctive tags on their bookmarks than Delicious users. Connotea users, when conducting their research or projects, may classify their online resources into specific topics for future retrieval or to facilitate division of labors, so that more unique tags are used to classify different topics within a project. Delicious users, in contrast, are likely to use more general tags to classify their bookmarks, such as music, sport, etc., in order to present a general topic.
Since no publication is found on studying the similarities and differences of users’ behavior in using two different bookmarking services that target at different user groups, this study helps fill the research gap by comparing the bookmarking and tagging behavior of Delicious and Connotea users.
5. Limitation and further research
This study examines the usage pattern in Connotea and Delicious, by analyzing the log data crawled from these two social bookmarking websites. The analysis can only identify the usage pattern and the difference in behavior of these two social bookmarking users, and the reason behind such pattern or difference is yet to be investigated. Further study can be done from the users’ perspective to examine in what ways social bookmarking users utilize the services and how effective the services are in helping them manage or organize their online resources.
6. Conclusion
Although social bookmarking has received growing attention and interest in the general public as well as in academia, the users’ bookmarking and tagging behavior have not been distinguished out of social bookmarking services which target at different user groups. The reasons behind such difference are yet to be discovered. The results of this study suggest that the bookmarking behaviors of Connotea and Delicious users have distinctive difference. Most of the Delicious users (in the sample pool) create bookmarks frequently, while Connotea users deviate a lot in their bookmarking behavior. The discrepancy in their bookmarking behavior can be partly explained by the existence period of user accounts, which has a greater effect on Connotea users then Delicious. The study also finds that there is a strong link between bookmarks created and tags used. In particular, Connotea users tend to use more unique tags for their bookmarks than Delicious users. Further investigation is needed to probe deeper into the reasons behind the social bookmarking users’ behavioral differences and assess the effectiveness of different bookmark management strategies employed by these users.
Acknowledgement
The research team would like to thank Mr. Wenfeng Han for his contribution in the preliminary analysis.
References
Al-Khalifa HS, Davis HC (2007): Towards Better Understanding of Folksonomic Patterns. In: Conference on Hypertext and Hypermedia, Proceedings of the 18th Conference on Hypertext and Hypermedia, Manchester, UK: ACM Press, pp.163–166.
Dye, J. (2006). Folksonomy: A game of high-tech (and high-stakes) tag. E-Content, 29(3), 38-43.
Golder, S.A., Huberman, B.A. (2006). Usage patterns of collaborative tagging systems. Journal of Information Science, 32(2), 198-208.
Gordon-Murnane, L. (2006). Social bookmarking, folksonomies and web 2.0 tools. Searcher, 14(6):26-38.
Hammond, T., Hannay, T., Lund, B., & Scott, J. (2005). Social bookmarking tools. A general review. Part 1. DLib Magazine, 12(1).
Hotho, A., Jäschke, R., Schmitz, C., Stumme, G.: Information retrieval in folksonomies: Search and ranking. In: Sure, Y., Domingue, J. (eds.) ESWC 2006. LNCS, vol. 4011, pp. 411–426. Springer, Heidelberg (2006)
Kipp, M.E.I. (2006a). @toread and cool: Tagging for time, task and emotion. 17th ASIS&T SIG/CR
Lund, B., Hammond, T., Flack, M., & Hannay, T. (2005). Social bookmarking tools (II). A case study - Connotea. D-Lib Magazine, 11(4).
Mathes, A. (2004). Folksonomies - Cooperative classification and communication through shared metadata. Computer Mediated Communication, LIS590CMC, 1-13.
Menchen, E. Feedback, Motivation and Collectivity in a Social Bookmarking System. In Kairosnews Computers and Writing Online Conference. 2005.
Millen, D., Feinberg, J., Kerr, B.: "Social Bookmarking in the Enterprise", ACM Queue, 3, 9 (2005), 28-35.
HELEN S. DU
Department of Computing, The Hong Kong Polytechnic University,
Hung Hom, Kowloon, Hong Kong
E-mail: cshelen@inet.polyu.edu.hk
SAMUEL K.W. CHU1 and FLORENCE T.Y. LAM2
Division of Information Technologies, Faulty of Education, The University of Hong Kong1
Department of Statistics and Actuarial Science, The University of Hong Kong2
E-mail: samchu@hkucc.hku.hk1 and florentia.lam@gmail.com2
Social bookmarking services have shown themselves as common and popular Internet tools by successfully acquiring millions of users, with Delicious being one of the most popular social bookmarking services to the public. While Delicious is used mainly for general purposes, Connotea, another social bookmarking site that primarily serves academic and scientific interests, has become equally popular among researcher groups. This paper attempts to analyze and compare users’ bookmarking and tagging behavior in Connotea and Delicious. The results show that there is a distinctive difference in usage behavior among these two groups of users. Delicious users create bookmarks more frequently than Connotea users, but Connotea users tend to use more distinctive tags for their bookmarks than Delicious users. Moreover, our result from the analysis indicates that the number of bookmarks created is a significant predictor of the quantity of tags used. This study is a starting point from which to explore the reasons behind the difference in social bookmarking and tagging behavior among different usage orientation groups.
1. Introduction
Social bookmarking sites have successfully acquired millions of online users in the recent years. A social bookmarking site provides the service for storing, sharing and discovering a collection of web bookmarks with the help of user-generated taxonomies (folksonomies). The term folksonomy came from the words taxonomy and folk, which is used to name the growing phenomenon of users collaboratively creating and managing metadata by tagging pieces of digital information with their own searchable keywords (Dye, 2006). Folksonomy is also known as social tagging. For simplicity, this paper uses the word tagging to refer to social tagging. In recent years, research has been done on tagging and folksonomy to understand folksonomic patterns (Al-Khalifa, & Davis, 2007) and their trends (Hotho, Jäschke, Schmitz, & Stumme 2006). Very few articles have been written on social bookmarking. The ones that have been published mainly give definitions on social bookmarking and its related concepts (Golder & Huberman, 2006) and provide general discussions on social bookmarking tools (Hammond, Hannay, Lund and Scott, 2005; Lund, Hammond, Flack and Hannay, 2005).
Delicious, created by Joshua Schachter in 2003, is one of the most popular social bookmarking service sites. Until now, it has acquired over 5 million users and 150 million bookmarked URLs [http://en.wikipedia.org/wiki/Delicious, viewed June 8, 2009]. Diigo is another similar social bookmarking service like Delicious, which allows users to bookmark and tag web pages. It also contains the functions to highlight and paste sticky notes, so that users can create their personal notes on the content of WebPages they visit. While Delicious and Diigo are social bookmarking tools that serve general purposes, some other social bookmarking tools such as CiteULike and Connotea are being targeted at academic and scientific purposes (Gordon-Murnane, 2006). These two social bookmarking service sites operate in the similar style as of Delicious, in addition to capturing the bibliographic information from scientific articles and journals.
From a preliminary analysis using a data mining technique (association rule mining), we find a notable difference in the usage pattern of the top 100 frequently used tags between Connotea and Delicious. As shown in Figure 1, Delicious has more diversified relationships of tags than Connotea.
Fig. 1. Association rules found in Connotea and Delicious with .1% minimum support and 50% confidence.
Note. Each element in the graph represents a tag. The arrow connecting any tags implies their association relationship. For example, one user uses the tag “protein” may also use “human” to bookmark the same URL (“unidirectional relationship”); one user uses the tag “protein” will also use “RNA” and vice versa (“bidirectional relationship”).
Based on the above preliminary analysis, we suspect that there exists different tagging behavior among the users of Delicious and Connotea that requires further investigation. This investigation is worthwhile as Delicious and Connotea are two very popular social bookmarking sites that serve different target user groups (general-purpose users and researchers respectively). So far, no published articles on social bookmarking have yet attempted to examine and compare social bookmarking services with different orientations. Therefore, this study investigates both Delicious and Connotea, and examines their similarities and differences in users’ bookmarking and tagging behavior.
2. Literature review
2.1 What is social bookmarking?
When people start to rely on the use of Internet, one of the greatest challenges they face is to remember and retrieve items that they have previously found to be useful. The common approach to arrange information on the Web is through the use of personal bookmarks (Millen, Feinberg, & Kerr, 2005). The desire to share information among communities has led to the development of shared bookmarking systems. Social bookmarking is a way to locate, classify and share Internet resources through the use of shared lists of user-created Internet bookmarks. Social bookmarking tools allow users to create tags for bookmarks they saved, and organize all users’ tags so that users can search and browse the tags to find out not only their own bookmarks but also other users’ bookmarks.
2.2 Social bookmarking services
Delicious [http://delicious.com], being the most commonly and widely used social bookmarking service among online users, is a server-based system with a simple-to-use interface that allows users to organize and share bookmarks on the Internet. Hotho et al. (2006) saw such service yields benefits for each individual user (e.g. organizing one’s bookmarks in a browser-independent, persistent fashion) without too much overhead, while there is a proliferation of resources in the Web that makes it difficult to remain up-to-date and to keep track on documents that are related to one’s own area of interest. The usage of social bookmarking services indicates that folksonomy-based approaches seem to be the solution to overcome this difficulty.
A review on social bookmarking services, particularly Connotea [www.connotea.org], and their advantages were summarized by Hammond et al. (2005) and in its companion paper by Lund et al. (2005). Connotea makes sharing among personal collections of resources much easier than before. Instead of placing materials hierarchically in folders, Connotea allows users to create simple tags to the bookmarks. Tagging allows the organization of bookmarks to be more flexible, multi-faceted and spacious. Furthermore, all bookmarks posted by these tools are visible to registered users, which take the concept of sharing to a higher level and benefit “not just from the ease with which it allows explicit sharing with friends and colleagues, but from many users storing their bookmarks in the same space” (Lund et al. 2005, p. 4).
2.3 Advantages/ Disadvantages of social bookmarking
Several researchers (Hotho et al, 2006; Laura, 2006; Menchen, 2005; Millen, 2006) have identified various advantages of social bookmarking. Social bookmarking services provide online storage that facilitates a single repository (Menchen, 2005). They allow an individual to create personal collections of bookmarks and facilitate sharing of bookmarks (Menchen, 2005; Millen, 2006). Social bookmarking services also allow users to create tags, which help organize and categorize users’ collection of information (Gordon-Murnane, 2006; Millen, 2006). By searching or browsing specific tags, users can retrieve all the items bookmarked by other users, which are tagged with specific keywords (Gordon-Murnane, 2006). Besides, bookmarks become portable since they are web-based. Users can access to the links and sites from computers, in contrast to the traditional way of accessing bookmarks from dedicated computers (Gordon-Murnane, 2006).
However, tags lack hierarchy so that a search of a specific term will only yield results on that term and not provide the full body of related terms that might be relevant to the user’s information needs and goals (Gordon-Murnane, 2006). Tags can be considered uncontrolled vocabulary shared across the entire social bookmarking system, and they have inherent ambiguity as different users apply terms to documents in different ways. Tags have no synonym control in the system, words with similar intended meanings, plural and singular forms will also appear in the system (Mathes, 2004). Al-Khalifa and Davis (2007) conducted a study to analyze tags in Delicious and classify them into three groups. They found that the non-standardized forms of tags make it difficult to capture them in a general form. For example, there are spelling variations, compound tags with different combination etc.
2.4 Users’ tagging behavior
Hotho et al. (2006) showed how topic-specific trends can be discovered in Delicious, one of the folksonomy-based systems. They collected users’ profiles and tags in Delicious within a time frame, and analyzed the tags to discover topic-specific trends. They argued that their analysis can be done regardless of the types of the underlying resources which would make folksonomies interesting for multimedia applications.
Kipp’s (2007) study found that a surprising number of tags used in the three social bookmarking tools examined (Delicious, Connotea and Citeulike) were not subject-related. These non subject-related tags can be classified into two groups: affective tags and time and tasks related tags. These behaviors have suggested that “users appear to want to store more than just the subject of the documents they are bookmarking” (Kipp, 2007, p. 3). They revealed specifically the users’ desire to “express an emotional connection to the document” and to “attach personal information management information to documents” (Kipp, 2008, p. 5).
Golder and Huberman (2005, p. 198) investigated the usage patterns of the tagging system in order to “identify regularities in user activity, tag frequencies, kinds of tags used, bursts of popularity in bookmarking and the stability in the relative proportions of tags within a given URL”. It is found that the frequencies of tagging vary much among users. In particular, there is no strong connection between the length of the user account’s existence and the number of days the user takes to create one or more bookmarks. Similarly, the number of bookmarks created by users has very little association with the number of tags used in each bookmark as well. However, a user’s tagging behavior could possibly be used to reflect his or her development of interests. For instance, as a tag grows steadily over time, it might indicate the user’s continual interest in that particular subject. On the other hand, if one tag suddenly grows rapidly, it might reveal the user’s newfound interest (Golder & Huberman, 2005, p. 202).
Other trends in social bookmarking are also observed, such as the time taken for an URL to reach its peak popularity. Although it is found that a majority of the URLs would indeed reach their peak of popularity as soon as they are introduced in Delicious, others have actually taken longer time before they are “rediscovered” and experience a sudden jump (Golder & Huberman, 2005, p. 204-205). In addition, it is often believed that social bookmarking is chaotic, unstructured, and imprecise because the collection of tags depends on each individual’s personal preference and level of knowledge. However, Golder and Huberman’s study claims that each tag’s frequency for a particular URL is a nearly fixed proportion of the total frequency of all tags used. More interestingly, this stability becomes apparent after fewer than 100 bookmarks (Golder & Huberman, 2005, p. 205-206).
2.5 Research Gap
Golder and Huberman (2006) analyzed the structure of collaborative tagging systems of Delicious and discovered regular patterns in user’s activity and tag frequencies. He expected that such findings could be applied to other similar tagging systems. However, little has been done regarding his claim. Therefore, it is worthwhile to conduct a research using similar methodology to that of Golder and Huberman (2006), but with a social bookmarking service designed for scientific purposes, such as Connotea.
The preliminary investigation using data mining also suggests that there is different usage pattern between the two sites – users of Delicious create tags for all kinds of topics while Connotea users create tags more for research purposes. Therefore, it is meaningful to study the users’ tagging behavior in Delicious alongside Connotea to find out if there are any significant similarities and differences between social bookmarking services of different orientations.
3. The Study
3.1 Data Collection and Sampling
Analysis is conducted on two sets of data (one for Delicious and one for Connotea) which contain the activity profiles of some of its users. A program is written using Java 1.4.2 to retrieve data that query both databases. The ways to decide which user should be included are slightly different between Delicious and Connotea because Connotea has a smaller user base than Delicious. For Delicious, the user’s history of activity profile is collected and analyzed if one or more of his/her bookmarks reached the top 500 most popular bookmarks on Feb 16, 2009 (Midnight GMT). This information is found through the public RSS feeds of the ‘popular’ page and then a portion of the website from the user’s activity profile is crawled. For Connotea, the list of users is compiled by examining at the bookmarks posted between April 1, 2008 (Midnight GMT) and May 31, 2008 (23:59 GMT).
This time period is chosen because of two reasons: 1) Most of the researchers and lecturers would have their summer break in June, it would be best to set the time frame before summer to represent the use of Connotea among researchers; 2) the user base for Connotea is smaller than Delicious and a longer period is needed to have a similar amount of bookmarks tagged. Based on the list of users who bookmarked in the aforementioned period, all the bookmarks they have posted since the activation of their accounts till Feb 17, 2009 (0:00:00 GMT) were retrieved. This procedure stopped when over 200,000 bookmarks from both databases were collected. As a result, we collected 5454 users’ data from Connotea and 440 users’ data from Delicious. Note here that the sample size of Connotea is over twelve times more than that of Delicious because Delicious users on average create much more bookmarks than Connotea users (see Table 1).
Table 1. Descriptive Statistics of user profile
Delicious Connotea
Mean 1404.175 49.724
Median 711.500 4
Min 3 1
Max 19100 15067
Note. The values are presented in number of bookmarks
3.2 Data Analysis
This analysis studies the bookmarking and tagging behavior of Connotea and Delicious respectively. The result between Connotea and Delicious is also compared to examine any significant differences.
3.2.1 The descriptive statistics of users’ bookmarking behavior
From the statistics (see Table 2), it is found that the bookmarking behaviors of Connotea and Delicious users differ significantly. It is shown from the standard deviation value and range value that users of Delicious do not deviate a lot in their bookmarking behavior. All users in the sample pool create at least one bookmark every month. However, some Connotea users may only create one bookmark every two years.
Table 2. Descriptive Statistics of users’ bookmarking behavior of Connotea and Delicious
Delicious Connotea
Mean 1.485 8.394
Standard Deviation 2.356 23.614
Minimum 0 0
Maximum 24.221 676.863
Note. The values are present in days.
3.2.2 Relationship between the length of the account since activation and the total number of bookmarks they created
In this regression analysis, both Connotea and Delicious’ data are best fitted into the exponential model. See Table 3 for the exponential regression results.
Delicious. As a result, the length of the user account since activation is a significant predictor of users’ bookmarking behavior variance (F¬ = 74.323, p < .05). In other words, the length of the user account since activation accounts for 14.3% of users’ bookmarking behavior variance. Therefore, it was likely for Delicious users who have created an account for a longer time to create more bookmarks (β = .381, p < .05).
Connotea. The length of the user account’s existence since activation is a significant predictor of users’ bookmarking behavior variance (F¬ = 5806.223, p < .05). In other words, the length of the user account’s since activation accounts for 51.6% of users’ bookmarking behavior variance. Connotea users who have created an account for a longer time were therefore more likely to create more bookmarks (β = .718, p < .05).
Table 3. Univariate Regressions: Delicious versus Connotea
(Dependent Variable: number of bookmarks created)
R2 df F Sig. β t
Delicious a .145 1 74.323 .000 .381 8.621
Connotea a .516 1 5806.223 .000 .718 76.199
a predictor: the length of the user account’s existence in days
In addition, a two-way between-subjects analysis of variance is conducted to compare the effect of the length of account since activiation on number of bookmarks created in Delicious and Connotea. The between-subjects factors are length of the existence of the account and nominal value with two levels (1 or 0 representing Connotea or Delicious, respectively). A significant interaction effect was found, F (1, 5890) = 1071.98, p < .000. The positive coefficient of interaction term (0.0039) suggests the length of account’s existence has a greater effect on number of bookmarks created in Connotea than Delicious.
3.3.3 Relationship between the number of bookmarks a user creates and the number of tags they use in those bookmarks.
In this analysis, the overall relationship between the number of bookmarks users created and the number of tags they used in those bookmarks will be analyzed. The lower end of the scale (users keeping fewer than 30 bookmarks) and the upper end (users keeping more than 500 bookmarks) will also be examined. Both Connotea and Delicious’ data are best fitted into the exponential model for regression analysis. Table 4 gives the exponential regression results.
Delicious. As a result, the number of bookmarks a Delicious user creates is a significant predictor of users’ tagging behavior variance (F¬ = 1207.6, p < .05). In other words, the number of bookmarks a user creates is able to account for 73.7% of users’ tagging behavior variance. Users who create more bookmarks are also likely to use more tags for those bookmarks (β = .858, p < .05). The relationship is found to be weaker at the lower end of the scale with users having fewer than 30 bookmarks (R = .652, R2 = .425, p < .05), but stronger at the upper end with users having more than 500 bookmarks (R = .716, R2 = .512, p < .05).
Connotea. The number of bookmarks a Connotea user creates is a significant predictor of users’ tagging behavior variance (F¬ = 38941.46, p < .05). The results suggest that the number of bookmarks a user creates is able to account for 87.7% of users’ tagging behavior variance. Users who create more bookmarks will also likely to use more tags for those bookmarks (β = .937, p < .05). It is stronger at the lower end of the scale, with users having fewer than 30 bookmarks (R = .845, R2 = .715, p < .05) but comparatively weaker at the upper end, with users having more than 500 bookmarks (R = .825, R2 = .681, p < .05).
Table 4. Univariate Regressions: Delicious Versus Connotea
(Dependent Variable: number of tags they use in bookmarks)
N R2 df F β t
Delicious a
Overall
442
.737
1
1207.6**
.858
34.751
Lower end (< 30 bookmarks) 14 .425 1 9.604** .652 3.099
Upper end (>500 bookmarks) 277 .512 1 289.76** .716 17.022
Connotea a
Overall
5455
.877
1
38941.466**
.937
197.336
Lower end (< 30 bookmarks)
Upper end (>500 bookmarks) 4334 .715 1 10848.069** .845 104.154
92 .681 1 194.527** .825 13.947
a predictor: number of bookmarks a user creates
**p < .01.
In addition, a two-way between-subjects analysis of variance is conducted to compare the effect of the number of bookmarks created on number of tags used in Delicious and Connotea. The between-subjects factors are number of bookmarks created and nominal value with two levels (1 and 0 represents Connotea and Delicious respectively). A significant interaction effect was found, F (1, 5884) = 5.79, p < .05. The negative coefficient of interaction term (-0.054) suggests the number of bookmarks created has a greater effect on number of tags used in Delicious than Connotea.
3.3.3 Proportion of unique tags over all tags used per user account
Figure 3 shows a significant difference in tagging behavior between Connotea and Delicious users. Almost half of Connotea users use mainly distinctive tags to organize or classify their bookmarks, while most of Delicious users use comparatively less unique tags for organization and classification.
Fig. 3. Proportion of unique tags over all tags used per user account in Connotea and Delicious
Note. 6 users from Delicious are excluded since they do not have any tags.
4. Discussion and Implications
Golder and Huberman (2006) expected that their findings on Delicious can be applied to other similar tagging systems. However, our findings suggest that Delicious and Connotea are quite different from one another. Sections 3.2.1 and 3.2.2 analyze the bookmarking behavior of Delicious and Connotea users, and show a distinctive difference between them. The reason behind this finding can be explained by the different user orientations of these two social bookmarking services. Since Delicious is used for general purposes, its users may use Delicious to bookmark websites related to personal interests or entertainment etc., which may occur frequently. Connotea, on the other hand, is a social bookmarking service that caters specifically to the management of scientific references. Users of Connotea mainly use its tools for doing research or for other academic purposes. Once they complete their assignments or research projects, they usually stop using Connotea until they start another project.
Sections 3.2.3 and 3.2.4 examine the tagging behavior of Delicious and Connotea users. Section 3.2.3 shows that both groups of users will use more tags if they have more bookmarks. It is noted that Connotea requires users to use at least one tag per bookmark while Delicious has no such restriction. These findings may suggest that tagging is a useful tool for users to manage, classify and organize bookmarks so that users are willing to create tags on bookmarks even if they are not required to do so. However, users of Delicious and Connotea have different tagging behavior on the bookmarks they created. Section 3.2.4 shows that Connotea users are more likely to use distinctive tags on their bookmarks than Delicious users. Connotea users, when conducting their research or projects, may classify their online resources into specific topics for future retrieval or to facilitate division of labors, so that more unique tags are used to classify different topics within a project. Delicious users, in contrast, are likely to use more general tags to classify their bookmarks, such as music, sport, etc., in order to present a general topic.
Since no publication is found on studying the similarities and differences of users’ behavior in using two different bookmarking services that target at different user groups, this study helps fill the research gap by comparing the bookmarking and tagging behavior of Delicious and Connotea users.
5. Limitation and further research
This study examines the usage pattern in Connotea and Delicious, by analyzing the log data crawled from these two social bookmarking websites. The analysis can only identify the usage pattern and the difference in behavior of these two social bookmarking users, and the reason behind such pattern or difference is yet to be investigated. Further study can be done from the users’ perspective to examine in what ways social bookmarking users utilize the services and how effective the services are in helping them manage or organize their online resources.
6. Conclusion
Although social bookmarking has received growing attention and interest in the general public as well as in academia, the users’ bookmarking and tagging behavior have not been distinguished out of social bookmarking services which target at different user groups. The reasons behind such difference are yet to be discovered. The results of this study suggest that the bookmarking behaviors of Connotea and Delicious users have distinctive difference. Most of the Delicious users (in the sample pool) create bookmarks frequently, while Connotea users deviate a lot in their bookmarking behavior. The discrepancy in their bookmarking behavior can be partly explained by the existence period of user accounts, which has a greater effect on Connotea users then Delicious. The study also finds that there is a strong link between bookmarks created and tags used. In particular, Connotea users tend to use more unique tags for their bookmarks than Delicious users. Further investigation is needed to probe deeper into the reasons behind the social bookmarking users’ behavioral differences and assess the effectiveness of different bookmark management strategies employed by these users.
Acknowledgement
The research team would like to thank Mr. Wenfeng Han for his contribution in the preliminary analysis.
References
Al-Khalifa HS, Davis HC (2007): Towards Better Understanding of Folksonomic Patterns. In: Conference on Hypertext and Hypermedia, Proceedings of the 18th Conference on Hypertext and Hypermedia, Manchester, UK: ACM Press, pp.163–166.
Dye, J. (2006). Folksonomy: A game of high-tech (and high-stakes) tag. E-Content, 29(3), 38-43.
Golder, S.A., Huberman, B.A. (2006). Usage patterns of collaborative tagging systems. Journal of Information Science, 32(2), 198-208.
Gordon-Murnane, L. (2006). Social bookmarking, folksonomies and web 2.0 tools. Searcher, 14(6):26-38.
Hammond, T., Hannay, T., Lund, B., & Scott, J. (2005). Social bookmarking tools. A general review. Part 1. DLib Magazine, 12(1).
Hotho, A., Jäschke, R., Schmitz, C., Stumme, G.: Information retrieval in folksonomies: Search and ranking. In: Sure, Y., Domingue, J. (eds.) ESWC 2006. LNCS, vol. 4011, pp. 411–426. Springer, Heidelberg (2006)
Kipp, M.E.I. (2006a). @toread and cool: Tagging for time, task and emotion. 17th ASIS&T SIG/CR
Lund, B., Hammond, T., Flack, M., & Hannay, T. (2005). Social bookmarking tools (II). A case study - Connotea. D-Lib Magazine, 11(4).
Mathes, A. (2004). Folksonomies - Cooperative classification and communication through shared metadata. Computer Mediated Communication, LIS590CMC, 1-13.
Menchen, E. Feedback, Motivation and Collectivity in a Social Bookmarking System. In Kairosnews Computers and Writing Online Conference. 2005.
Millen, D., Feinberg, J., Kerr, B.: "Social Bookmarking in the Enterprise", ACM Queue, 3, 9 (2005), 28-35.
PATTERNS OF SEMANTIC-BASED
PATTERNS OF SEMANTIC-BASED KNOWLEDGE CLASSIFICATION, ORGANIZATIONAND ACQUISITION PROCESSES FOR PRODUCT DESIGN ENTERPRISES
YINGLIN WANG; JIANMEI GUO;
.Department of Computer Science and Engineering, Shanghai Jiao Tong University, Shanghai 200240, China, { ylwang, guojianmei}@sjtu.edu.cn
XIJUAN LIU
Mechanical School, Shanghai Dianji University, Shanghai 200240
Former knowledge engineering researches aims at boosting automatic reasoning, however recent knowledge management researches focus on promoting the knowledge sharing and reusing among the peoples. Because of the different aims between the two directions, former knowledge representations schemas, such as rule based representation, frame from knowledge engineering researches are not fit for the current knowledge management scenarios. In this paper, for the purpose of building knowledge management systems for product design enterprises, knowledge items are classified into seven types based on the semantics of their usage. Then their representations are discussed respectively. Based on the above classifications, a knowledge representation meta-model and a basic domain ontology reference model for cooperative knowledge management systems are put forward. The reference model is an abstraction that can be reused and extended in knowledge management systems of different enterprises. Finally, the patterns of knowledge acquisition processes in cooperative knowledge management scenarios of product design processes are studied.
1. Introduction
Knowledge representation and acquisition have been studied for many years, and a lot of methods have been put forward. The most common and famous methods are rules, semantic networks, frames and case based methods (Feigenbaum 1984; Bruce 1984; Randall 1986; Bergmann 2002; Bueno 2005). For instance, semantic networks are used to represent knowledge. Each node represents a concept and arcs are used to define relations between the concepts. Frame has its own name and a set of attributes, or slots which contain values. However, the former methods usually aim at supporting automatic reasoning, rather than helping human beings in their design decisions. The purpose is ambitious, but it is beyond the reach, just because knowledge in its nature is so complex that the process may not be fully automated. E.g., fulfilling product design tasks needs very complex cognitive, creative and synthetic activities, which are far beyond what the automatic reasoning can deal with. Up to now, complex product design tasks still mainly rely on the efforts of human beings. Therefore representations of knowledge should be suitable for helping human beings in their routine works, so that the knowledge discovered can be easily represented, well-organized, easily searched when required.
Hence, in order to capture and reuse knowledge in business processes to improve the product, processes and services, the concept of Knowledge Management (KM) was put forward since the 1991 (Nonaka 1991). KM comprises a range of practices used in an organization to identify, create, represent, distribute and enable adoption of insights and experiences, which are either embodied in individuals or embedded in organizational processes or practice. As a consequence, many large companies and non-profit organizations have resources dedicated to internal KM efforts (Addicott, McGivern & Ferlie 2006). Several consulting companies also exist that provide strategy and advice regarding KM to these organizations.
Relating to knowledge representation, different frameworks for distinguishing between knowledge were studied. One proposed framework distinguishes knowledge between tacit knowledge and explicit knowledge (Polanyi 1966). Tacit knowledge represents internalized knowledge which is hard to be expressed by artificial language and may not be consciously aware of how he or she accomplishes particular tasks, and this is different from explicit knowledge, which represents knowledge that the individual holds consciously in mental focus, in a form that can easily be communicated to others (Alavi & Leidner 2001). Early research suggested that a successful KM effort needs to convert internalized tacit knowledge into explicit knowledge in order to share it, but the same effort must also permit individuals to internalized and make personally meaningful any codified knowledge retrieved from the KM effort. Nonaka proposed a model (SECI for Socialization, Externalization, Combination, Internalization) which considers a spiraling knowledge process interaction between explicit knowledge and tacit knowledge (Nonaka & Takeuchi 1995). In this model, knowledge follows a cycle in which implicit knowledge is 'extracted' to become explicit knowledge, and explicit knowledge is "re-internalized" into implicit knowledge.
At the meantime, researchers put forward frameworks from other different perspectives. Davenport believed that knowledge may or may not be clearly expressed, taught, used during observation; may be detailed or outlined; may be complex or simple, undocumented and documented (Davenport 1993). Specifically, in the product design field. Ropohl divided design knowledge into 4 types from the angle of systems philosophy: technical know-how, functional rules, structural rules, and socio-technological understanding (Ropohl 1997). Bayazit defined 4 types of design knowledge on the basis of design methodology: procedural knowledge, declarative knowledge, design normative knowledge and collaborative design knowledge (Bayazit 1993). He considered collaborative design knowledge as knowledge about team working.
The existing classification frameworks of knowledge are helpful, but some of them fail to clarify knowledge in a clear and systematic way, and the dependency relationships of different kinds of knowledge are missing. Besides, the semantics of the usages, e.g., the connections between knowledge, specialties and tasks in a domain, have not been emphasized. This makes knowledge hard to be maintained and reused. Hence, effective way of knowledge representation and organization for knowledge management activities should be studied further.
Besides the knowledge classification frameworks, the key issue in knowledge management is knowledge acquisition, which has been studied extensively, and many approaches have been proposed. In general sense, knowledge acquisition contains three types: discovery by personal, discovery through cooperation in an organization, and discovery via data mining. From the cognitive view point, Nonaka proposed a paradigm for managing the dynamic aspects of organizational knowledge creating processes (Nonaka 1994). Its central theme is that organizational knowledge is created through a continuous dialogue between tacit and explicit knowledge. He also identifies four patterns of interaction involving tacit and explicit knowledge: socialization, externalization, internalization, and combination. These patterns form a spiral model for organizational knowledge creation. Related to Nonaka's framework, recent years, the researchers studied the patterns of Collaborative Knowledge Management (CKM) and Personal Knowledge Management (PKM) processes respectively. CKM is based on a software environment where people work on-line; they continuously contribute to the collective knowledge, which is then made available to everyone. The key idea is to motivate participation in a collective knowledge creation process by supporting online environments for collaborative work, and to harvest the value of collaborative work using high precision search, alerting, and related KM technologies. Typical researches of CKM are Ramesh 's and Kathryn's. Ramesh identified problems associated with knowledge management in the context of new product development by cross-functional collaborative teams. A prototype system that captures and manage tacit and explicit process knowledge is also discussed (Ramesh 1999). Kathryn introduced the concept of knowledge management for product innovation and presents a collaborative knowledge management tool specifically designed to help manage a portfolio of product innovation projects in a distributed environment (Kathryn 2003). Lihui present a review of existing research, projects, and applications in the domain of collaborative conceptual design, based on the Internet and Web technologies. The purpose of the review is to understand the needs for conceptual engineering design, to clarify the current conceptual design practice, to classify the available technologies, and to study the future trend in this area (Lihui 2002). Compared with CKM, PKM is focused on personal productivity improvement for knowledge workers in their working environments. While the focus is the individual, the goal of PKM is to enable individuals to operate better both within the formal structure of organizations and in looser work groupings. This is as different from KM as traditionally viewed, which appears to be focused on enabling the corporation to be more effective by "recording" and making available what its workers know. Wright studied how individual workers apply knowledge processes to support their day-to-day work activities, and presented an emergent model that links distinctive types of problem solving activities with specific cognitive, information, social and learning competencies, supported by individual, social and organizational enablers (Wright 2005). Jefferson holds that PKM is aimed at helping the individual to overcome the frustrations associated with information overload and allowing them to improve their personal effectiveness (Theresa 2005). The third approach of knowledge acquisition is acquiring knowledge via data mining or using machine learning techniques, such as decision tree, association rules, and neural network. Rubin, S.H. views knowledge mining as an extension of data mining (Rubin 1997). Qingzhang proposed a framework named Intelligent Knowledge Discovery System (IKDS) which help users to select appropriate data mining algorithms to discover useful knowledge (Qingzhang 2007).
Existing researches of knowledge acquisition present the basic schemes of knowledge acquisition process, but patterns and mechanism for integration between them are still rare. Therefore, how different kinds of knowledge acquisition processes and business processes are integrated, and how it is implemented, need to be reconsidered in the complex context of business processes. Besides, in product design domains, knowledge and data might be very complex, much of which can not be represented using structural description. As a result of this, full-automatic knowledge acquisition is hard to be used. So, in which way the automatic mining tools are used need to be discussed further.
Based on the existing researches and the above assumption, we classify the knowledge items used in product design routines into seven abstract types: concepts, relations, rules, methods, processes, know-where, and instances according to the semantics of the roles and usages of them. Based on this classification and the domain ontology, the knowledge items can be organized in a systematic way. Besides, we propose an integration model of knowledge acquisition processes, which combines personal knowledge acquisition processes, collaborative knowledge acquisition processes, and the data mining processes together, forming a seamless lifecycle knowledge management model. The implementation aspects are also discussed in which we use ontology as the basis for building the system.
2. Knowledge Types in Complex Product Knowledge Management Scenarios
Based on the existing results mentioned in section 1, and the empirical experiences, we define seven knowledge types for complex product knowledge management, the semantics of each type are as follows.
Concepts are the basic units of knowledge which are the basics for understanding and communications between different persons and organizations. According to the semantic theory of concepts, concepts are abstract objects (Eric et al, 2007). Many philosophers consider concepts to be a fundamental ontological category of being. From the viewpoint of artificial intelligence domains, concepts are considered as subsets of objects. From the engineering domain, concepts may be physical notations, such as temperature, speed, power etc. They may be special shapes on a component, such as fillet, chamfer etc. They may be special components or mechanical devices, such as motorcycle, gear-box, etc. They may be abstract notations, such as changes, conflicts etc. For represent design knowledge, concepts should be represented and defined first. Important concepts, such as tasks, products of enterprises should be defined. Concepts are the key elements of a domain ontology.
Relations reflect the possible dependencies and associations between a set of variables or objects. Relations can be represented as formula, relational tables, graphs or figures which may represent the qualitative dependencies between variables. Relations can be distinguished into two types: those which always are true for all the instantiation of related variables, and those are true for only specific values of the variables. The former kind can be called formulas in a general sense, while the latter kind can be called individual relations. E.g., the fact "some employee x has participated in a project y" is an individual relation. But the assertion "any person who has been responsible for a job of kind x is able to do the jobs of same kind" should be considered a formula.
Rules are guidelines of what is and what is not allowed in the design decision process. For example, a design taboo is a kind of rules. Some of them may be deduced from corresponding relations. When choosing some parameters when others has been known in a design problem, the related formula relations may be used.
Methods could be directly used to solve certain problems, for instance, the method to increase the rigidity of materials, a way to reduce the energy consumption of a specific machine, or a technique to split the mesh when using the finite element software.
Processes include a set of dependent activities in order to fulfill special tasks. Processes may allow iterations. For example, diagnosis of design problems follows a specific process. The product design is a collaborative process in which variant of resources are involved.
Know-where is the special knowledge of where to find the knowledge or persons to solve problems at hand. For example, knowledge map can be thought as this kind of knowledge.
Figure 1 The knowledge representation meta-model
Instances are specific results or experiences of solving problems, including instances of processes and the results. Instances are the original resources through which the abstract concepts, relations, rules, methods, processes and know-where knowledge can be discovered and generalized. For example, a piece of design case of an artifact, in which a set of design concepts, values and relations between design parameters, the related design methods, and the persons who designed it are described.
The relationship of the seven types of knowledge is shown in Figure 1. In this figure, concepts are the basis for all other kinds of the knowledge representation; relations will be defined on the top of concepts; rules may be derived through relations; the methods, know-where knowledge, and processes are depicted on top of concepts and relations. Those kinds of knowledge can be abstracted from instances; or in the reverse direction, the instances can be instantiated from the more abstracted types. A unified ontology meta-model is used as the framework to describe the domain ontology.
Figure 2 The basic domain ontology reference model
The above kinds of knowledge may be represented in structural forms or represented in natural language forms. Here the guideline is that only when automatic reasoning will have great benefits to the users, should the knowledge be transformed into structural forms. When there is a structural form of a knowledge item, the corresponding natural language form will still be stored in the knowledge base, which is used for human beings to understand it.
Based on the above general classifications of knowledge, we put forward basic domain ontology model as an initialization for building cooperative knowledge management systems (see Figure 2). The model includes the essential classes and relations, such as the classes of organizations, employees, specialties and products, concept descriptions etc. For instance, a unit in an organization includes a group of employees who cooperate in their work to achieve certain common goals of designing, manufacturing, warehousing, and selling the final product. Based on the basic domain ontology model, more specific concepts and relations can be developed.
3. Knowledge Acquisition Process Integration Model
The above classification and the reference knowledge meta-model provide the basis for the acquisition of knowledge. However, knowledge acquisition in cooperative environments is far more than trivial pursuit. Through the analysis of knowledge discovery, we seamlessly integrated three types of knowledge acquisition patterns, i.e., personal discovery, discovering through cooperation, discovering through data mining, into a unified models of the knowledge acquisition processes. In the first pattern, the employee accepts task assigned to he or her; retrieve related knowledge from the enterprise knowledge base (or the knowledge might be pushed to the employee automatically according the task type, etc.); execute the task accordingly and obtain results; through the results, the background knowledge and the former results he or she already knows, the employee independently finds knowledge, record it down and submit it after further modification and edit using the terminology and format of the ontology. Software modules can be embedded in the personal work platforms for editing and submitting knowledge at anytime. In the second pattern, knowledge is discovered by several persons who cooperatively work together towards a final solution for complex tasks. In this kind of knowledge acquisition pattern, a workflow process might be involved and the responsibility of the participants during the process is clearly defined. In Figure 2, for simplicity, we do not distinguish the above two patterns although there are slight differences between them. For both kinds of patterns, after the submission of knowledge, an intermediate state named quasi-knowledge is given to it. A knowledge item of quasi-knowledge will be published on the bulletin board of the enterprise and wait other persons to recommended if they found it valuable and worthy of recommendation. After the recommendation, an evaluation process will follow to judge the correct and the value of the knowledge. Then, the evaluation work will be assigned to related experts, according to some business rules. One of the possible rule might be "IF the employee x work in the specialty y, and the knowledge z to be evaluated involves the specialty y, then x can be assigned the evaluation task of knowledge z".
The third pattern of data mining techniques is integrated in the model in the following ways. When new results are added into the database of the enterprise during execution of a task or several tasks, rules might be generalized through data mining, but as the rules might be nonsense, so they must be checked by the workers before they can be submitted for further evaluation. The automatic content analysis techniques, e.g., text similarity analysis, can also be used to check whether the new findings are similar as the knowledge that already exists in the enterprise's knowledge base to avoid further duplicate efforts of the evaluation for it; or the contradictions might be find through automatic content analysis, so the old knowledge might need to be revised. The automatic content analysis techniques can also be used in the collaborative evaluation process thereafter in the similar ways (see Figure 3). In the evaluation process, the similar knowledge of the evaluated knowledge in the knowledge base will be found and pushed to the reviewers for references to avoid the duplicate efforts of evaluation and the possible redundancy of knowledge in the knowledge base.
Figure 3 Knowledge Acquisition Process Integration Model
4. Case Study
4.1. The organization of knowledge
Domain ontology (which is built on top of the reference model in Figure 2) plays a central role in the above integration model. Each object in the system must be represented via the terminology of concepts and relations of the domain ontology and the standard language, such that it can be understood by others; the representation via ontology will enable some possible automation of the business and knowledge processes. Besides, the system will be more flexible via using domain ontology schema.
For example, as a concept of mechanical engineering, "shaft bearing" may be defined with a general description of its function, and the parameter's (attributes) description as well as some graphics that shows the meaning of them (see Figure 4).
For instance, part of the definition of the "shaft bearing" can be as follows:
Section 1: Basic Definition
Definition of shaft bearing: A device that supports, guides, and reduces the friction of motion between fixed shaft and moving machine parts.
The super-class: Mechanical Connections;
The figure of shaft bearing: see Figure 4.
Section 2: Attributes of a Shaft Bearing;
Attribute (Parameter) 1: Name: Load of the Shaft, W; The semantics: the total force that the bearing is designed to withstand. See also the definition of Mechanical Load; Value types: numerical; Unit: lbf;
Attribute (Parameter) 2: Name: Diameter of the Shaft, d; The semantics: the allowed diameter of the shaft the bearing can support; Value types: numerical; Unit: inch;
… …
Section 3: Relations between the Parameters of a Shaft Bearing;
1) The general description:
The initial step in the selection and sizing of a bearing involves determination of the operation bearing pressure. Bearing pressure is defined as the load divided by the projected area:
2) The formula:
P = 4.4482 x W / (d x L ) where:
P = Bearing pressure, MPa
W = Load, N
d = Shaft diameter, mm
L = Bearing length, mm
3) Variants of the formula;
This formula gives the average pressure in MPa, that the bearing supports. Elevated temperature reduces load capacity; lower temperature generally increases static load capacity.
Section 4: Rules
Title: Rules for selecting bearing proportions;
Content: Optimum performance can be achieved by specifying a length to inside diameter ratio (L/d) ranging from 0.5 to 2.0. Values of L/d less than 1.0 result in easier escape for wear debris and less sensitivity to shaft deflection and misalignment. There may also be some cost advantage in using a bearing with a small L/d ratio. If the L/d is higher than 2.0, distortions or misalignment may cause stress concentrations and excessive localized heating. So make sure don't let the L/d >2.0.
Section 5: Methods
Title: Method when a long bearing is required.
Content: When a long bearing is required, it is advisable to consider using two bearings with a small gap between them or to increase the inside diameter, d, and re-estimate the bearing geometry.
Section 6: Instances
Based on the above definition of "shaft bearing" in the domain ontology, the instances of the concept can be specified thereafter. The graph is helpful to explain the meaning of the relationships of parameters. There may be many formats that can represent the instances of a class (concept). One way is using tables . Each row in the table means a kind of valid combination of these parameters. In product design domain, many of the knowledge take this kind of form.
In this approach, the elements defined in section 1 and section 2 are the basis for all the other parts. For example, when a new bearing instance is available and need to be input to the knowledge base, then the definition in section 1 and section 2 should be used by the system to generate a user interface for the user to input the values of the parameters of the new instance.
4.2. Integrated Knowledge Acquisition and Evolution Processes
Now let's discuss a case in which former recognized knowledge need to be amended. If a designer uses the knowledge of shaft bearing of the knowledge base in a new scenario, and find that the rule or method is invalid or insufficient for this new situation, then the person may need to amend the rule or method of the knowledge base. E.g., if he or she uses a shaft bearing whose L/d ratio < 2.0, but it results in a failure afterwards, then the former rule should be modified to reflect the failure. The contradiction with the former rule may be found by the automatic analysis tools, or some data mining tools may be used, to aid the person to generalize the rules.
Then the person may put forward a new version of the rule, edit and then submit it as a quasi knowledge to be evaluated. For choosing the right experts to evaluate the new findings, the specialty of the knowledge might be indicated and used. E.g., assuming that a piece of knowledge “K_1” mainly relates to a specialty “S_1”, the specialty “S_1” is associated with some employee instances through the “has-Specialty” relation of the class “Employee” in the domain ontology. If these related employees are available, they can be arranged to review the new knowledge; otherwise, the employees from neighboring specialties are selected. Such neighboring specialties can be found in terms of the specialty hierarchy or the similarity of them. If the new knowledge involves different specialties, the employees from the different specialties will be selected.
The knowledge-specialty-employee relationship, used in the above knowledge acquisition process, is defined in ontology meta-model and is instantiated in object-instance layer. Thus, there are three layers: ontology meta-model, domain ontology, and object instances. The rules of employee assignment are business rules, which can be represented on top of the concepts and relations in domain ontology. Similarly, the specialty hierarchy, being a part of domain ontology, can also be used to routine business processes. Take task assignment for example, some task can be assigned to some employees or departments according to the relationship between task types and specialty types.
Acknowledgements
The paper was supported by the National High-tech Research and Development Project of China (863) under the grant No. 2009AA04Z106 and the National Science Foundation of China (NSFC) under the grant No. 60773088.
5. Conclusion
According to the practical requirements of building collaborative knowledge management systems in product design enterprises, we classify the knowledge items into seven types based on the semantics of their usage. In this model, all the elements of knowledge can be specified based on an ontology meta-model. The concept and relation are the basis of others. The rules, methods and processes can be defined using the terminologies of the concept and relation layers. As the dependencies between the elements of knowledge are clearly described, hence knowledge can be systematically maintained. We also propose an integration model that seamlessly integrates three types of knowledge acquisition processes and business processes based on the domain ontology. The pattern of the knowledge representations and acquisition we summarized in this paper are general abstractions from real applications, thus it can be used as a reference model to build the future collaborative knowledge management systems in enterprises, especially for complex product development domains.
References
Addicott, Rachael; Gerry McGivern & Ewan Ferlie (2006), "Networks, Organizational Learning and Knowledge Management: NHS Cancer Networks", Public Money & Management 26 (2): 87-94
Alavi, Maryam & Dorothy E. Leidner (2001), "Review: Knowledge Management and Knowledge Management Systems: Conceptual Foundations and Research Issues", MIS Quarterly 25 (1): 107-136
Bayazit, N. (1993). Designing: design knowledge, design research, related sciences in M J de Vries, N Cross and D P Grant (eds). Design methodology and relationships with science, Kluwer Academic Publishers, Dordrecht.
Bergmann, R. (2002). "Experience management: foundations, development Methodology, and Internet-based applications". Berlin, Springer.
Bruce G. Buchanan, Edward H. Shortliffe (1984), "Rule Based Expert Systems: The MYCIN Experiments of the Stanford Heuristic Programming Project", Addison-Wesley, Reading, MA.
Bueno,T. C. D'Agostini (2005). "Knowledge Engineering Suite: A tool to create ontologies for automatic knowledge representation in knowledge-based systems", Lecture Notes in Computer Science. Electronic Government: 4th International Conference, EGOV 2005.
Davenport, T. H.,L. Prusak (1993). Working Knowledge:How Organization Manage What They Know. Boston, Havard Business School Press.
Eric Margolis and Stephen Laurence (2007) "The Ontology of Concepts: Are Concepts Abstract Objects or Mental Representations?", Nous, 41(4) 561-93.
Feigenbaum, E. A. (1984), "Knowledge Engineering: The Applied Side of Artificial Intelligence.", Annals New York Academy of Sciences, 91-107.
Kathryn Cormican, David O'Sullivan.” A collaborative knowledge management tool for product innovation management”, International Journal of Technology Management, 26(1):53-67, 2003
Lihui Wang, Weiming Shen, Helen Xie, Joseph Neelamkavil and Ajit Pardasani. (2002). Collaborative conceptual design—state of the art and future trends. Computer-Aided Design Volume 34, Issue 13, November 2002, Pages 981-996
M. Polanyi. (1966). “The Tacit Dimension”. Routledge & Kegan Paul, London.
Nonaka. (1994) “A Dynamic Theory of Organizational Knowledge Creation”, Organization Science, 5 (1) 14-37.
Nonaka, Ikujiro (1991), "The knowledge creating company", Harvard Business Review 69 (6 Nov-Dec): 96-104
Nonaka, Ikujiro; Takeuchi, Hirotaka (1995). The knowledge creating company: how Japanese companies create the dynamics of innovation. New York: Oxford University Press. pp. 284.
Qingzhang Chen; Chao Chen; Xiaoying Chen; (2007). An Intelligent Knowledge Discovery System with a Novel Knowledge Acquisition Methodology.
Ramesh, B.,A. Tiwana (1999). "Supporting Collaborative Process Knowledge management in New Product Development Teams." Decision Support Systems 27: 213-235.
Randall Davis (1986), "Knowledge-Based Systems", Science, 28 February, 231(4741):957 - 963
Ropohl, G. (1997). "Knowledge types in technology." International Journal of Technology and Design Education 7(1-2): 65-72.
Rubin, S.H. (1997). "Data vs knowledge mining a crossing of theories", 1997 IEEE International Conference on Systems, Man, and Cybernetics, 1997.
Theresa L. Jefferson. (2005). “Taking it personally: personal knowledge management”, VINE, 36(1): 35-37.
Wright, Kirby. (2005). "Personal knowledge management: supporting individual knowledge worker performance", Knowledge Management Research and Practice 3: 156–165.
YINGLIN WANG; JIANMEI GUO;
.Department of Computer Science and Engineering, Shanghai Jiao Tong University, Shanghai 200240, China, { ylwang, guojianmei}@sjtu.edu.cn
XIJUAN LIU
Mechanical School, Shanghai Dianji University, Shanghai 200240
Former knowledge engineering researches aims at boosting automatic reasoning, however recent knowledge management researches focus on promoting the knowledge sharing and reusing among the peoples. Because of the different aims between the two directions, former knowledge representations schemas, such as rule based representation, frame from knowledge engineering researches are not fit for the current knowledge management scenarios. In this paper, for the purpose of building knowledge management systems for product design enterprises, knowledge items are classified into seven types based on the semantics of their usage. Then their representations are discussed respectively. Based on the above classifications, a knowledge representation meta-model and a basic domain ontology reference model for cooperative knowledge management systems are put forward. The reference model is an abstraction that can be reused and extended in knowledge management systems of different enterprises. Finally, the patterns of knowledge acquisition processes in cooperative knowledge management scenarios of product design processes are studied.
1. Introduction
Knowledge representation and acquisition have been studied for many years, and a lot of methods have been put forward. The most common and famous methods are rules, semantic networks, frames and case based methods (Feigenbaum 1984; Bruce 1984; Randall 1986; Bergmann 2002; Bueno 2005). For instance, semantic networks are used to represent knowledge. Each node represents a concept and arcs are used to define relations between the concepts. Frame has its own name and a set of attributes, or slots which contain values. However, the former methods usually aim at supporting automatic reasoning, rather than helping human beings in their design decisions. The purpose is ambitious, but it is beyond the reach, just because knowledge in its nature is so complex that the process may not be fully automated. E.g., fulfilling product design tasks needs very complex cognitive, creative and synthetic activities, which are far beyond what the automatic reasoning can deal with. Up to now, complex product design tasks still mainly rely on the efforts of human beings. Therefore representations of knowledge should be suitable for helping human beings in their routine works, so that the knowledge discovered can be easily represented, well-organized, easily searched when required.
Hence, in order to capture and reuse knowledge in business processes to improve the product, processes and services, the concept of Knowledge Management (KM) was put forward since the 1991 (Nonaka 1991). KM comprises a range of practices used in an organization to identify, create, represent, distribute and enable adoption of insights and experiences, which are either embodied in individuals or embedded in organizational processes or practice. As a consequence, many large companies and non-profit organizations have resources dedicated to internal KM efforts (Addicott, McGivern & Ferlie 2006). Several consulting companies also exist that provide strategy and advice regarding KM to these organizations.
Relating to knowledge representation, different frameworks for distinguishing between knowledge were studied. One proposed framework distinguishes knowledge between tacit knowledge and explicit knowledge (Polanyi 1966). Tacit knowledge represents internalized knowledge which is hard to be expressed by artificial language and may not be consciously aware of how he or she accomplishes particular tasks, and this is different from explicit knowledge, which represents knowledge that the individual holds consciously in mental focus, in a form that can easily be communicated to others (Alavi & Leidner 2001). Early research suggested that a successful KM effort needs to convert internalized tacit knowledge into explicit knowledge in order to share it, but the same effort must also permit individuals to internalized and make personally meaningful any codified knowledge retrieved from the KM effort. Nonaka proposed a model (SECI for Socialization, Externalization, Combination, Internalization) which considers a spiraling knowledge process interaction between explicit knowledge and tacit knowledge (Nonaka & Takeuchi 1995). In this model, knowledge follows a cycle in which implicit knowledge is 'extracted' to become explicit knowledge, and explicit knowledge is "re-internalized" into implicit knowledge.
At the meantime, researchers put forward frameworks from other different perspectives. Davenport believed that knowledge may or may not be clearly expressed, taught, used during observation; may be detailed or outlined; may be complex or simple, undocumented and documented (Davenport 1993). Specifically, in the product design field. Ropohl divided design knowledge into 4 types from the angle of systems philosophy: technical know-how, functional rules, structural rules, and socio-technological understanding (Ropohl 1997). Bayazit defined 4 types of design knowledge on the basis of design methodology: procedural knowledge, declarative knowledge, design normative knowledge and collaborative design knowledge (Bayazit 1993). He considered collaborative design knowledge as knowledge about team working.
The existing classification frameworks of knowledge are helpful, but some of them fail to clarify knowledge in a clear and systematic way, and the dependency relationships of different kinds of knowledge are missing. Besides, the semantics of the usages, e.g., the connections between knowledge, specialties and tasks in a domain, have not been emphasized. This makes knowledge hard to be maintained and reused. Hence, effective way of knowledge representation and organization for knowledge management activities should be studied further.
Besides the knowledge classification frameworks, the key issue in knowledge management is knowledge acquisition, which has been studied extensively, and many approaches have been proposed. In general sense, knowledge acquisition contains three types: discovery by personal, discovery through cooperation in an organization, and discovery via data mining. From the cognitive view point, Nonaka proposed a paradigm for managing the dynamic aspects of organizational knowledge creating processes (Nonaka 1994). Its central theme is that organizational knowledge is created through a continuous dialogue between tacit and explicit knowledge. He also identifies four patterns of interaction involving tacit and explicit knowledge: socialization, externalization, internalization, and combination. These patterns form a spiral model for organizational knowledge creation. Related to Nonaka's framework, recent years, the researchers studied the patterns of Collaborative Knowledge Management (CKM) and Personal Knowledge Management (PKM) processes respectively. CKM is based on a software environment where people work on-line; they continuously contribute to the collective knowledge, which is then made available to everyone. The key idea is to motivate participation in a collective knowledge creation process by supporting online environments for collaborative work, and to harvest the value of collaborative work using high precision search, alerting, and related KM technologies. Typical researches of CKM are Ramesh 's and Kathryn's. Ramesh identified problems associated with knowledge management in the context of new product development by cross-functional collaborative teams. A prototype system that captures and manage tacit and explicit process knowledge is also discussed (Ramesh 1999). Kathryn introduced the concept of knowledge management for product innovation and presents a collaborative knowledge management tool specifically designed to help manage a portfolio of product innovation projects in a distributed environment (Kathryn 2003). Lihui present a review of existing research, projects, and applications in the domain of collaborative conceptual design, based on the Internet and Web technologies. The purpose of the review is to understand the needs for conceptual engineering design, to clarify the current conceptual design practice, to classify the available technologies, and to study the future trend in this area (Lihui 2002). Compared with CKM, PKM is focused on personal productivity improvement for knowledge workers in their working environments. While the focus is the individual, the goal of PKM is to enable individuals to operate better both within the formal structure of organizations and in looser work groupings. This is as different from KM as traditionally viewed, which appears to be focused on enabling the corporation to be more effective by "recording" and making available what its workers know. Wright studied how individual workers apply knowledge processes to support their day-to-day work activities, and presented an emergent model that links distinctive types of problem solving activities with specific cognitive, information, social and learning competencies, supported by individual, social and organizational enablers (Wright 2005). Jefferson holds that PKM is aimed at helping the individual to overcome the frustrations associated with information overload and allowing them to improve their personal effectiveness (Theresa 2005). The third approach of knowledge acquisition is acquiring knowledge via data mining or using machine learning techniques, such as decision tree, association rules, and neural network. Rubin, S.H. views knowledge mining as an extension of data mining (Rubin 1997). Qingzhang proposed a framework named Intelligent Knowledge Discovery System (IKDS) which help users to select appropriate data mining algorithms to discover useful knowledge (Qingzhang 2007).
Existing researches of knowledge acquisition present the basic schemes of knowledge acquisition process, but patterns and mechanism for integration between them are still rare. Therefore, how different kinds of knowledge acquisition processes and business processes are integrated, and how it is implemented, need to be reconsidered in the complex context of business processes. Besides, in product design domains, knowledge and data might be very complex, much of which can not be represented using structural description. As a result of this, full-automatic knowledge acquisition is hard to be used. So, in which way the automatic mining tools are used need to be discussed further.
Based on the existing researches and the above assumption, we classify the knowledge items used in product design routines into seven abstract types: concepts, relations, rules, methods, processes, know-where, and instances according to the semantics of the roles and usages of them. Based on this classification and the domain ontology, the knowledge items can be organized in a systematic way. Besides, we propose an integration model of knowledge acquisition processes, which combines personal knowledge acquisition processes, collaborative knowledge acquisition processes, and the data mining processes together, forming a seamless lifecycle knowledge management model. The implementation aspects are also discussed in which we use ontology as the basis for building the system.
2. Knowledge Types in Complex Product Knowledge Management Scenarios
Based on the existing results mentioned in section 1, and the empirical experiences, we define seven knowledge types for complex product knowledge management, the semantics of each type are as follows.
Concepts are the basic units of knowledge which are the basics for understanding and communications between different persons and organizations. According to the semantic theory of concepts, concepts are abstract objects (Eric et al, 2007). Many philosophers consider concepts to be a fundamental ontological category of being. From the viewpoint of artificial intelligence domains, concepts are considered as subsets of objects. From the engineering domain, concepts may be physical notations, such as temperature, speed, power etc. They may be special shapes on a component, such as fillet, chamfer etc. They may be special components or mechanical devices, such as motorcycle, gear-box, etc. They may be abstract notations, such as changes, conflicts etc. For represent design knowledge, concepts should be represented and defined first. Important concepts, such as tasks, products of enterprises should be defined. Concepts are the key elements of a domain ontology.
Relations reflect the possible dependencies and associations between a set of variables or objects. Relations can be represented as formula, relational tables, graphs or figures which may represent the qualitative dependencies between variables. Relations can be distinguished into two types: those which always are true for all the instantiation of related variables, and those are true for only specific values of the variables. The former kind can be called formulas in a general sense, while the latter kind can be called individual relations. E.g., the fact "some employee x has participated in a project y" is an individual relation. But the assertion "any person who has been responsible for a job of kind x is able to do the jobs of same kind" should be considered a formula.
Rules are guidelines of what is and what is not allowed in the design decision process. For example, a design taboo is a kind of rules. Some of them may be deduced from corresponding relations. When choosing some parameters when others has been known in a design problem, the related formula relations may be used.
Methods could be directly used to solve certain problems, for instance, the method to increase the rigidity of materials, a way to reduce the energy consumption of a specific machine, or a technique to split the mesh when using the finite element software.
Processes include a set of dependent activities in order to fulfill special tasks. Processes may allow iterations. For example, diagnosis of design problems follows a specific process. The product design is a collaborative process in which variant of resources are involved.
Know-where is the special knowledge of where to find the knowledge or persons to solve problems at hand. For example, knowledge map can be thought as this kind of knowledge.
Figure 1 The knowledge representation meta-model
Instances are specific results or experiences of solving problems, including instances of processes and the results. Instances are the original resources through which the abstract concepts, relations, rules, methods, processes and know-where knowledge can be discovered and generalized. For example, a piece of design case of an artifact, in which a set of design concepts, values and relations between design parameters, the related design methods, and the persons who designed it are described.
The relationship of the seven types of knowledge is shown in Figure 1. In this figure, concepts are the basis for all other kinds of the knowledge representation; relations will be defined on the top of concepts; rules may be derived through relations; the methods, know-where knowledge, and processes are depicted on top of concepts and relations. Those kinds of knowledge can be abstracted from instances; or in the reverse direction, the instances can be instantiated from the more abstracted types. A unified ontology meta-model is used as the framework to describe the domain ontology.
Figure 2 The basic domain ontology reference model
The above kinds of knowledge may be represented in structural forms or represented in natural language forms. Here the guideline is that only when automatic reasoning will have great benefits to the users, should the knowledge be transformed into structural forms. When there is a structural form of a knowledge item, the corresponding natural language form will still be stored in the knowledge base, which is used for human beings to understand it.
Based on the above general classifications of knowledge, we put forward basic domain ontology model as an initialization for building cooperative knowledge management systems (see Figure 2). The model includes the essential classes and relations, such as the classes of organizations, employees, specialties and products, concept descriptions etc. For instance, a unit in an organization includes a group of employees who cooperate in their work to achieve certain common goals of designing, manufacturing, warehousing, and selling the final product. Based on the basic domain ontology model, more specific concepts and relations can be developed.
3. Knowledge Acquisition Process Integration Model
The above classification and the reference knowledge meta-model provide the basis for the acquisition of knowledge. However, knowledge acquisition in cooperative environments is far more than trivial pursuit. Through the analysis of knowledge discovery, we seamlessly integrated three types of knowledge acquisition patterns, i.e., personal discovery, discovering through cooperation, discovering through data mining, into a unified models of the knowledge acquisition processes. In the first pattern, the employee accepts task assigned to he or her; retrieve related knowledge from the enterprise knowledge base (or the knowledge might be pushed to the employee automatically according the task type, etc.); execute the task accordingly and obtain results; through the results, the background knowledge and the former results he or she already knows, the employee independently finds knowledge, record it down and submit it after further modification and edit using the terminology and format of the ontology. Software modules can be embedded in the personal work platforms for editing and submitting knowledge at anytime. In the second pattern, knowledge is discovered by several persons who cooperatively work together towards a final solution for complex tasks. In this kind of knowledge acquisition pattern, a workflow process might be involved and the responsibility of the participants during the process is clearly defined. In Figure 2, for simplicity, we do not distinguish the above two patterns although there are slight differences between them. For both kinds of patterns, after the submission of knowledge, an intermediate state named quasi-knowledge is given to it. A knowledge item of quasi-knowledge will be published on the bulletin board of the enterprise and wait other persons to recommended if they found it valuable and worthy of recommendation. After the recommendation, an evaluation process will follow to judge the correct and the value of the knowledge. Then, the evaluation work will be assigned to related experts, according to some business rules. One of the possible rule might be "IF the employee x work in the specialty y, and the knowledge z to be evaluated involves the specialty y, then x can be assigned the evaluation task of knowledge z".
The third pattern of data mining techniques is integrated in the model in the following ways. When new results are added into the database of the enterprise during execution of a task or several tasks, rules might be generalized through data mining, but as the rules might be nonsense, so they must be checked by the workers before they can be submitted for further evaluation. The automatic content analysis techniques, e.g., text similarity analysis, can also be used to check whether the new findings are similar as the knowledge that already exists in the enterprise's knowledge base to avoid further duplicate efforts of the evaluation for it; or the contradictions might be find through automatic content analysis, so the old knowledge might need to be revised. The automatic content analysis techniques can also be used in the collaborative evaluation process thereafter in the similar ways (see Figure 3). In the evaluation process, the similar knowledge of the evaluated knowledge in the knowledge base will be found and pushed to the reviewers for references to avoid the duplicate efforts of evaluation and the possible redundancy of knowledge in the knowledge base.
Figure 3 Knowledge Acquisition Process Integration Model
4. Case Study
4.1. The organization of knowledge
Domain ontology (which is built on top of the reference model in Figure 2) plays a central role in the above integration model. Each object in the system must be represented via the terminology of concepts and relations of the domain ontology and the standard language, such that it can be understood by others; the representation via ontology will enable some possible automation of the business and knowledge processes. Besides, the system will be more flexible via using domain ontology schema.
For example, as a concept of mechanical engineering, "shaft bearing" may be defined with a general description of its function, and the parameter's (attributes) description as well as some graphics that shows the meaning of them (see Figure 4).
For instance, part of the definition of the "shaft bearing" can be as follows:
Section 1: Basic Definition
Definition of shaft bearing: A device that supports, guides, and reduces the friction of motion between fixed shaft and moving machine parts.
The super-class: Mechanical Connections;
The figure of shaft bearing: see Figure 4.
Section 2: Attributes of a Shaft Bearing;
Attribute (Parameter) 1: Name: Load of the Shaft, W; The semantics: the total force that the bearing is designed to withstand. See also the definition of Mechanical Load; Value types: numerical; Unit: lbf;
Attribute (Parameter) 2: Name: Diameter of the Shaft, d; The semantics: the allowed diameter of the shaft the bearing can support; Value types: numerical; Unit: inch;
… …
Section 3: Relations between the Parameters of a Shaft Bearing;
1) The general description:
The initial step in the selection and sizing of a bearing involves determination of the operation bearing pressure. Bearing pressure is defined as the load divided by the projected area:
2) The formula:
P = 4.4482 x W / (d x L ) where:
P = Bearing pressure, MPa
W = Load, N
d = Shaft diameter, mm
L = Bearing length, mm
3) Variants of the formula;
This formula gives the average pressure in MPa, that the bearing supports. Elevated temperature reduces load capacity; lower temperature generally increases static load capacity.
Section 4: Rules
Title: Rules for selecting bearing proportions;
Content: Optimum performance can be achieved by specifying a length to inside diameter ratio (L/d) ranging from 0.5 to 2.0. Values of L/d less than 1.0 result in easier escape for wear debris and less sensitivity to shaft deflection and misalignment. There may also be some cost advantage in using a bearing with a small L/d ratio. If the L/d is higher than 2.0, distortions or misalignment may cause stress concentrations and excessive localized heating. So make sure don't let the L/d >2.0.
Section 5: Methods
Title: Method when a long bearing is required.
Content: When a long bearing is required, it is advisable to consider using two bearings with a small gap between them or to increase the inside diameter, d, and re-estimate the bearing geometry.
Section 6: Instances
Based on the above definition of "shaft bearing" in the domain ontology, the instances of the concept can be specified thereafter. The graph is helpful to explain the meaning of the relationships of parameters. There may be many formats that can represent the instances of a class (concept). One way is using tables . Each row in the table means a kind of valid combination of these parameters. In product design domain, many of the knowledge take this kind of form.
In this approach, the elements defined in section 1 and section 2 are the basis for all the other parts. For example, when a new bearing instance is available and need to be input to the knowledge base, then the definition in section 1 and section 2 should be used by the system to generate a user interface for the user to input the values of the parameters of the new instance.
4.2. Integrated Knowledge Acquisition and Evolution Processes
Now let's discuss a case in which former recognized knowledge need to be amended. If a designer uses the knowledge of shaft bearing of the knowledge base in a new scenario, and find that the rule or method is invalid or insufficient for this new situation, then the person may need to amend the rule or method of the knowledge base. E.g., if he or she uses a shaft bearing whose L/d ratio < 2.0, but it results in a failure afterwards, then the former rule should be modified to reflect the failure. The contradiction with the former rule may be found by the automatic analysis tools, or some data mining tools may be used, to aid the person to generalize the rules.
Then the person may put forward a new version of the rule, edit and then submit it as a quasi knowledge to be evaluated. For choosing the right experts to evaluate the new findings, the specialty of the knowledge might be indicated and used. E.g., assuming that a piece of knowledge “K_1” mainly relates to a specialty “S_1”, the specialty “S_1” is associated with some employee instances through the “has-Specialty” relation of the class “Employee” in the domain ontology. If these related employees are available, they can be arranged to review the new knowledge; otherwise, the employees from neighboring specialties are selected. Such neighboring specialties can be found in terms of the specialty hierarchy or the similarity of them. If the new knowledge involves different specialties, the employees from the different specialties will be selected.
The knowledge-specialty-employee relationship, used in the above knowledge acquisition process, is defined in ontology meta-model and is instantiated in object-instance layer. Thus, there are three layers: ontology meta-model, domain ontology, and object instances. The rules of employee assignment are business rules, which can be represented on top of the concepts and relations in domain ontology. Similarly, the specialty hierarchy, being a part of domain ontology, can also be used to routine business processes. Take task assignment for example, some task can be assigned to some employees or departments according to the relationship between task types and specialty types.
Acknowledgements
The paper was supported by the National High-tech Research and Development Project of China (863) under the grant No. 2009AA04Z106 and the National Science Foundation of China (NSFC) under the grant No. 60773088.
5. Conclusion
According to the practical requirements of building collaborative knowledge management systems in product design enterprises, we classify the knowledge items into seven types based on the semantics of their usage. In this model, all the elements of knowledge can be specified based on an ontology meta-model. The concept and relation are the basis of others. The rules, methods and processes can be defined using the terminologies of the concept and relation layers. As the dependencies between the elements of knowledge are clearly described, hence knowledge can be systematically maintained. We also propose an integration model that seamlessly integrates three types of knowledge acquisition processes and business processes based on the domain ontology. The pattern of the knowledge representations and acquisition we summarized in this paper are general abstractions from real applications, thus it can be used as a reference model to build the future collaborative knowledge management systems in enterprises, especially for complex product development domains.
References
Addicott, Rachael; Gerry McGivern & Ewan Ferlie (2006), "Networks, Organizational Learning and Knowledge Management: NHS Cancer Networks", Public Money & Management 26 (2): 87-94
Alavi, Maryam & Dorothy E. Leidner (2001), "Review: Knowledge Management and Knowledge Management Systems: Conceptual Foundations and Research Issues", MIS Quarterly 25 (1): 107-136
Bayazit, N. (1993). Designing: design knowledge, design research, related sciences in M J de Vries, N Cross and D P Grant (eds). Design methodology and relationships with science, Kluwer Academic Publishers, Dordrecht.
Bergmann, R. (2002). "Experience management: foundations, development Methodology, and Internet-based applications". Berlin, Springer.
Bruce G. Buchanan, Edward H. Shortliffe (1984), "Rule Based Expert Systems: The MYCIN Experiments of the Stanford Heuristic Programming Project", Addison-Wesley, Reading, MA.
Bueno,T. C. D'Agostini (2005). "Knowledge Engineering Suite: A tool to create ontologies for automatic knowledge representation in knowledge-based systems", Lecture Notes in Computer Science. Electronic Government: 4th International Conference, EGOV 2005.
Davenport, T. H.,L. Prusak (1993). Working Knowledge:How Organization Manage What They Know. Boston, Havard Business School Press.
Eric Margolis and Stephen Laurence (2007) "The Ontology of Concepts: Are Concepts Abstract Objects or Mental Representations?", Nous, 41(4) 561-93.
Feigenbaum, E. A. (1984), "Knowledge Engineering: The Applied Side of Artificial Intelligence.", Annals New York Academy of Sciences, 91-107.
Kathryn Cormican, David O'Sullivan.” A collaborative knowledge management tool for product innovation management”, International Journal of Technology Management, 26(1):53-67, 2003
Lihui Wang, Weiming Shen, Helen Xie, Joseph Neelamkavil and Ajit Pardasani. (2002). Collaborative conceptual design—state of the art and future trends. Computer-Aided Design Volume 34, Issue 13, November 2002, Pages 981-996
M. Polanyi. (1966). “The Tacit Dimension”. Routledge & Kegan Paul, London.
Nonaka. (1994) “A Dynamic Theory of Organizational Knowledge Creation”, Organization Science, 5 (1) 14-37.
Nonaka, Ikujiro (1991), "The knowledge creating company", Harvard Business Review 69 (6 Nov-Dec): 96-104
Nonaka, Ikujiro; Takeuchi, Hirotaka (1995). The knowledge creating company: how Japanese companies create the dynamics of innovation. New York: Oxford University Press. pp. 284.
Qingzhang Chen; Chao Chen; Xiaoying Chen; (2007). An Intelligent Knowledge Discovery System with a Novel Knowledge Acquisition Methodology.
Ramesh, B.,A. Tiwana (1999). "Supporting Collaborative Process Knowledge management in New Product Development Teams." Decision Support Systems 27: 213-235.
Randall Davis (1986), "Knowledge-Based Systems", Science, 28 February, 231(4741):957 - 963
Ropohl, G. (1997). "Knowledge types in technology." International Journal of Technology and Design Education 7(1-2): 65-72.
Rubin, S.H. (1997). "Data vs knowledge mining a crossing of theories", 1997 IEEE International Conference on Systems, Man, and Cybernetics, 1997.
Theresa L. Jefferson. (2005). “Taking it personally: personal knowledge management”, VINE, 36(1): 35-37.
Wright, Kirby. (2005). "Personal knowledge management: supporting individual knowledge worker performance", Knowledge Management Research and Practice 3: 156–165.
MOTIVATION, IDENTITY, AND AUTHORING
MOTIVATION, IDENTITY, AND AUTHORING OF THE WIKIPEDIAN
JOSEPH C. SHIH
Department of Information Management, Lunghwa University of Science and Technology, Taoyuan County, Taiwan, (R.O.C.)
Email: joseph@mail.lhu.edu.tw
C. K. FARN
Department of Information Management, National Central University,
Taoyuan County, Taiwan, (R.O.C.)
Email: ckfarn@mgt.ncu.edu.tw
Wikipedia is an online free encyclopedia which is edit by million people spontaneously. This article aims at how Wikipedians’ intrinsic and extrinsic motivations influence their volitional authoring. Research model posits that altruism, expected reputation, and expected money reward affect authoring behavior. More specific, this relationship is also mediated by both attitude and identity. We also regard perceived behavioral control as a critical role for fostering volitional authoring. Sample data were from “Wikipedian Discussion Board” of a famous BBS in Taiwan. All respondents (156 samples) had posted articles in the BBS, but not all had experienced authoring in Wikipedia. Structural equation modeling was used to test the research model. According to the result, we have insight into the motivation of Wikpedian’s volitional authoring.
1. Introduction
Wikipedia, a free online encyclopedia that anyone can edit, has a tremendous impact on how a great many writers gather information about the world (http://www.widipedia.org/). With no paid editors and written by numerous volunteers, Wikipedia is now emerged as the No. 1 go-to information source in the world. Wikipedia also now ranks eighth (July 2009) on the list of most visited sites on the Internet (http://alexa.com/), containing over 2.9 million articles in the English version (July 2009).
However, there are still a few people detracting the value of content (Badke, 2008), but the Wikipedian have persisted in pouring more and more items into the online encyclopedia. The phenomenon is interesting that there are still thousands of people participating in authoring items spontaneously, disregarding the controversial open source project being discredited by non-supporters. Furthermore, the economic exchange perspective posits that an individual’s decision making was found upon rational rule, as benefit surpassing cost. It looks as if the Wikipedian are not consistent with this tenet obviously. This paper, accordingly, is to understand how the Wikipedian’s motivation links to their volitional authoring, more specific, to examine their authoring behavior through the lens of attitude and virtual community identity.
2. Conceptual Background
2.1 Motivations for Participating in Open Source Project
Motivations are commonly categorized into extrinsic and intrinsic by researchers. Intrinsic motivations contain inherent satisfactions rather than their substantive consequence, such as volunteering and enjoying helping others which are congruent with one’s value system, (Kankanhalli et al., 2005; Ryan and Deci, 2000; Wasko and Faraj, 2005). On the other hand, extrinsic motivations means a focus on expected benefits of donating, where the extrinsic rewards are believed to exceed the contribution’s costs (Kankanhalli et al., 2005; Lerner and Tirole, 2002). Motivations of volunteering behavior have noticed in previous research (Clary et al., 1998), for example, Clary et al. (1998) classified six motivational categories for the volunteer. The categories include values, social, understanding, career, protective, and enhancement, and, contrasting with the functions, we describe the Wikipedian’s conceivable motivation in Table 1.
Table 1 Descriptions of Volunteer’s Motivations
Function Conceptual definition Description for Wikipedian
Values The individual volunteers in order to express or act on important values like humanitarianism. The Wikipedian feel it is important to help other by means of authoring.
Understanding The volunteer is seeking to learn more about the world or exercise skills that are often unused. Authoring lets the Wikipedian learn through direct and hands-on activities.
Enhancement One can grow and develop psychologically through volunteer activities. Authoring makes the Wikipedian feel better about themselves.
Career The volunteer has the goal of gaining career-related experience through volunteering. Authoring can help the Wikipedian to get experience related to current job.
Social Volunteering allows an individual to strengthen his or her social relationships. The Wikipedian identify to Wikipedia community.
Protective The individual uses volunteering to reduce negative feelings, such as guilt, or to address personal problems. Authoring is a good escape from day-to-day worry.
2.2 Attitude and Behavioral Control and Online Authoring
Theory of Planned Behavior (TPB) hypothesized that intention to perform a behavior is based on: attitudes, subjective norms and perceived behavioral control (Ajzen, 1991). Attitude towards performing the behavior is defined as a person’s general feeling of performing that behavior if a favorable or unfavorable action. Perceived behavioral control is assumed to reflect past experience as well as anticipated obstacles. The more opportunities and resource that individuals think they possess and the fewer obstacles they anticipate, the greater their perceived control over the behavior. Participating in online activities may be determined by attitude toward and perceived controllability over target behavior, such as blogging, using instant message software, or sharing knowledge in virtual community (Hsu and Chiu, 2008; Kuo and Young, 2008).
2.3 Social Identity and Online Prosocial Behavior
Social identity (SI) captures the main aspects of the individual’s identification with the group in the sense that a participant comes to view himself or herself as “belonging” to a certain group. A person professes “belonging” to a specific group is a psychological state, distinct from being a unique and distinct individual, conferring a collective representation of who one is (Hogg and Abrams, 1988). Many of previous studies related to online volitional behavior have emphasized the significance of virtual community identity. That participants consider themselves as parts of the target online group may strengthen the self-defining relation to the virtual community as well as foster prosocial behavior such like donating knowledge (Chiu et al., 2006; Jian and Jeffres, 2006; Ma and Agarwal, 2007).
3. Research Model and Hypotheses
3.1 Research Model
The present study posits intrinsic and extrinsic motivation associate with both attitude and virtual community identity (VC identity); and attitude, VC identity, and perceived behavior control (PBC) jointly affect authoring behavior. Beside, we also posits that PBC moderates both relationship of attitude-authoring and VC identity-authoring.
3.2 Research Hypotheses
Expressing or acting on important values like humanitarianism (Clary et al., 1998), it may be the most critical motivation of the Wikipedian. People engage in contribution of open content for their own sakes rather than for some external consequences (Hars and Ou, 2002). Positive relation between intrinsic motivation and authoring in Wikipedia was reported in previous study (Nov and Kuk, 2008). Altruism is a variant of intrinsic motivation in which a Wikipedia seeks to increase the welfare of others. Furthermore, altruism have widely held to be associated with positive attitude toward on-line helping behaviors, such as authoring in blog and participating in open source project (Hars and Ou, 2002; Hsu and Lin, 2008). For example, value of altruism, to which the volunteer may reflect their willingness to help people, is the primary function of the Wikipedian’s motivation (Clary and Snyder, 1999). Although Nov and Kuk (2008) proposed general intrinsic motivation in their study, more specific, we argue that altruism influences one’s attitude toward authoring in Wikipedia.
H1: Altruism positively associates with attitude toward authoring.
According to social categorization, people use demographic differences or distinguishable characteristic to categorize one another (Chatman and Spataro, 2005). Wikipedians, obviously different from non-participator, are willing to self-identify themselves as a specific collective, so that those in-group members emerge virtual community identity. We have two reasons to infer the relationship. One reason is that the value system of those who are higher extent of altruism may be more possible to align with the mission of Wikipedia. Altruistic people conceive the principle of unselfish concern for or devotion to the welfare of others, whereas writing items in Wikipedia premises on participant’s volitional behavior. This logic was also argued by Van Dick et al. (2006) that social identification is positively related to prosocial behavior. Another reason is that professional identity is an individual’s self-definition as a member of profession collective (Chreim et al., 2007). They emphasized “role identity” which can make that a professional conducts his or her work look more like professional. Therefore, those high altruistic people are happy to profess themselves as members of Wikipedian. Depending on the virtual community identity, not only a Wikipedian’s professional identity is acknowledged but also his or her altruistic authoring is encouraged. Thus, we propose the following hypothesis.
H2: Altruism positively associates with VC identity.
As economic considerations still critical to the volunteer, helper’s behaviors in cyberspace are as well, Donath (1999) remarked altruism is insufficient to explain helper’s motivation. This viewpoint is even more assertive by Kollock (1999) observing that motivation to contribute to online communities could spring from a variety of sources but none of them depended on altruism. In terms of their viewpoints, the volitional behaviors of the Wikipedian may be influenced by some economic considerations.
Earning reputation is regarded as an extrinsic motivation for a Wikipedian. Reputation is still an important asset, not only in real world but also in cyberspace (Wasko and Faraj, 2005), whereby an individual can leverage to achieve and maintain status within a collective. For the rank of Wikipedian growing with the number of articles edited and accumulated fame as a source of authority, cyberspatial reputation can be a strong incentive to engage authoring activities (Ciffolilli, 2003).
Next, we hypothesize the relationship between money reward and authoring behavior. According to economic exchange theory, one will make decisions by rational self-interest. Thus, prosocial behaviors such as knowledge sharing will occur if its rewards exceed its costs (Bock and Kim, 2002). In light of the rationale, the Wikipedian should have expected money reward while contributing knowledge to Wikipedia. Obviously, the Wikipedian did not conform to the rule. That the Wikipedian plainly understand no money reward but continually author leads to hypothesize the negative relation between money reward and attitude toward authoring. However, it doesn’t mean that Wikipedians hate money but merely they do not expect this extrinsic reward-in terms of money return-while authoring in Wikipedia.
H3: Extrinsic rewards associate with attitude toward authoring
H3a: Expected reputation positively associates with attitude toward authoring.
H3b: Expected money rewards negatively associates with attitude toward authoring.
For earning future reputation in virtual community, a Wikipedian regards social interaction with other aficionados as critical as authoring in virtual community. In additional to engage in authoring and contributing knowledge, the Wikipedian have to maintain a positive social relationship with other Wikipedians. The rationales are that a Wikipedian identifies to his or her community may satisfy the member’s need of self-defining and manifest the status of referent power (Bagozzi and Dholakia, 2002). Earning reputation not only is an antecedent for virtual community identity but also a driven force for authoring. Again, no money return is supposed to authoring in Wikipedia that obtaining identity from the community may, so that we hypothesize a negative relationship between expected money rewards and virtual community identity.
H4: Extrinsic rewards associate with VC identity.
H4a: Expect reputation positively associates with VC identity.
H4b: Expected money rewards negatively associates with VC identity.
Spending time and effort to complete articles, the Wikipedian consider authoring is consistent with their values and positive for online readers. Theory of Reasoned Action posits that behavioral formation is determined by attitude toward the behavior and subject norms of that behavior. The Wikipedian are inducing the authoring behavior while they have a positive attitude toward authoring. Previous studies related to online knowledge sharing have shown the consistent argument (Bock et al., 2005). For example, Kuo and Young (2008) found that the more favorable the individual’s attitude toward knowledge sharing practices, the stronger his/her intention to share knowledge in a teacher’s virtual community.
H5: Attitude toward authoring positively associates with authoring in Wikipedia.
Social identity model of deindividuation effects (SIDE) can demonstrate the relationship between Wikipedian’s community identity and authoring behavior. Volunteering allows an individual to strengthen his or her social relationships (Clay et al., 1998). In light of the rationale, community identity is useful in explaining individuals’ willingness to maintain committed relationship with the group (Nahapiet and Ghoshal, 1998), such as through sharing knowledge within the virtual community (Chiu et al., 2006). Empirical studies have reported that sense of belonging is important and has been used as a test for the presence of an online community. Jian and Jeffres (2006) argues that people are motivated to contribute to shared electronic databases because by doing so they will maintain and affirm relevant identities. Hsu and Lin (2008) argued the influence of social identification for blog users needs the perception of belonging to the virtual community. Likewise, the Wikipedian will present themselves as belonging to the community through continual authoring. Once being in Wikipedian community that suppresses individual differences and emphasizes a common Wikipedian membership, individuals have a high level of group identity and act according to the objective set up by the group (Kim, 2009). Herein, the objective set by Wikipedian community is contributing knowledge, i.e. authoring in Wikipedia.
H6: VC identity positively associates with authoring in Wikipedia.
As a Wikipedian believes himself or herself having sufficient resources to author in the community, the Wikipedian can complete the authoring behavior. Some obstacles such as no available time or incapable of computer skill make authoring impossible. Perceived behavioral control is an important antecedent of behavior theoretically (Ajzen, 1991) that the belief is a form of self-evaluation which influences decisions about what behaviors to undertake (Bandura, 1977). Previous empirical studies of online behavior also have confirmed the notion (Hsu and Chiu, 2004; Kuo and Young, 2008). Accordingly, we propose the following hypothesis.
H7: Perceived behavioral control positively associates with authoring in Wikipedia.
When the Wikipedian are highly perceived control over getting through with new items for Wikipedia, the positive attitude may foster the Wikipedian more authoring behavior. In term of contrast viewpoint-low perceived behavioral control, if a person thinks it good to author in Wikipedia but he does not have time to do this or he does not have enough computer skill to fill out the job, the person will decline the extent of attitude and then impede the authoring behavior. Likewise, when high perceived behavioral control, people with high VC identity will more possibly engaged in authoring in Wikipedia. Thus, we propose the hypothesis of moderating effect.
H8: The relationship between attitude and authoring in Wikipedia is moderated by perceived behavioral control (PBC). More specific, the relationship between attitude and authoring in Wikipedia is stronger when PBC is high than low.
H9: The relationship between VC identity and authoring in Wikipedia is moderated by perceived behavioral control (PBC). More specific, the relationship between VC identity and authoring in Wikipedia is stronger when PBC is high than low.
4. Research Methodology
A questionnaire was deployed on a web site in which allows respondents to answer the questionnaire. Structural equation modeling (SEM) technique was employed to test research model as well as H1 to H7; H8 and H9, the moderating effect of perceived behavioral control on attitude and identity to authoring behavior, were tested by hierarchical regression analysis. Followings are details of the current research methodology.
4.1 Scale Development
The survey questionnaire was designed on the basis of a comprehensive literature review and was refined via several runs of pretests and revisions.
Altruism was measured by a seven-point scale adopted from Chattopadhyay (1999), developed to capture a respondent’s seeking to increase the welfare of others. A sample item is “I will help co-worker (or classmate) who overloads with job (or school-homework).” Expected reputation was measure by a seven-point scale adopted from Constant et al. (1994). Attitude was measured by a seven-point scale adopted from Bock et al. (2005). Virtual community identity was measured by a seven-point scale adopted from Ellemers et al. (1999), developed to capture a respondent’s identification with the virtual community in the sense that the one comes to view himself or herself as a member of Wikipedia community. Perceived behavioral control was measured by a seven-point scale adopted from Armitage et al. (1999), developed to capture a respondent’s ease or difficulty of authoring in Wikipedia. After examining the nature of this scale, we regarded PBC as a formative construct that we would aggregate the score at consequent stage. In order to measure the extent in which the authors contribute to Wikipedia, we operationalized authoring in Wikipedia with two items, one is time consuming per week and another is the frequency of authoring per week. Detail items are omitted due to the limit of space in conference version.
4.2 Data Collection
A web-based site was deployed that respondents could visit to answer the questionnaire which was designed for collecting empirical data of the current study. To have a broad representation of both Wikipedian-authors and non-Wikipedian authors, participants were invited from the Wikipedian’s discussion board of PTT forum which is a famous bulletin board system in Taiwan. We invited them to participate in the survey via an email, in which attaching the web-site’s hyperlink, so that they could visit our web page to answer survey questions. They were also informed that we would donate 5 dollars to Wikipedian Foundation while a questionnaire was finished.
There were totally 181 respondents participating in the survey that the valid responding rate is 86% (156 valid) due to dropping 25 invalid questionnaires. Of 156 samples, the characteristics are demonstrated in Table 2
Table 2 Characteristics of Sample
Gender Education
Male 123 79% Under high school 13 8%
Female 33 21% High school 25 16%
University 80 51%
Authoring experience Graduated school 37 24%
Yes 107 69% PhD 1 1%
No 49 31%
Age
Years of using Wikipedia Under 15 7 5%
Under 1 37 24% 15 to 19 14 9%
1 73 47% 20 to 24 77 49%
2 32 21% 25 to 29 40 26%
3 9 6% 30 to 34 13 8%
4 4 3% 35 to 40 3 2%
5 and above 1 1% 40 above 2 1%
5. Data Analysis
5.1 Test of Measurement Model
Initial results of the CFA indicates that model were not fit the data well. A careful and iterative inspection of LISREL output revealed that some items did not load on the designated latent factors appropriately, such as standardized loading < 0.6 or associated with high modification indices. All indices are above cut-off value except GFI is slightly lower. The detail results are demonstrated in Table 3.
Before testing the structural model, it is also necessary to examine whether the measurement model had a satisfactory level of validity and reliability. While CR is greater than 0.7 and AVE is greater than 0.5, it implies that the variance captured by the latent construct is more than that by error component (Bagozzi et al., 1991). That is, each measure is accounting for 50 percent or more of the variance of the underlying latent variable (Chin, 1998). As the reports in Table 3, CRs and AVEs are all above recommended cut-off values that the scale is of internal consistency reliability.
Convergent validity ensures that all items measure a single latent construct, and it is established if all item loadings are greater than or equal to the recommended cut-off level of 0.70 (Bassellier et al., 2003). Our results showed that almost loadings of each latent variable are above the cut-off value that only two items’ are slightly lower. The details are also exhibited in Table 3.
Discriminant validity reflects the level to which the measures for each dimension are distinctively different from each other. We applied the chi-square difference test to assess the discriminant validity of the measurement model (Bassellier et al., 2003). Accordingly, we conducted 15 pair-wise tests (six constructs) that the results are reported in Table 4. All Δχ2 differences are significant above the level of Pr[χ2(1) ≥ 3.84]=0.05, indicating strong support for discriminant validity (Bassellier et al., 2003; Venkatraman, 1989). Additionally, the correlation matrixes are reported in Table 5.
Table 3 Results of Measurement Model Test
Factor Composition Reliability AVE Items Loading t-value Error
Altruism .84 .57 AL1 .63 8.67 .60
AL2 .71 10.17 .49
AL3 .92 14.28 .16
AL4 .73 10.43 .47
Expected Reputation .88 .64 ER1 .76 11.26 .43
ER2 .91 14.86 .17
ER3 .88 14.01 .23
ER4 .63 8.92 .60
Expected Money Reward --- --- EMR --- --- ---
Attitude .91 .71 AT1 .78 11.92 .39
AT2 .83 12.89 .32
AT3 .89 14.39 .22
AT4 .87 13.87 .25
Virtual Community Identity .83 .55 CI1 .72 10.09 .49
CI2 .74 10.58 .45
CI3 .74 10.60 .45
CI4 .76 11.00 .42
Perceived Behavioral Control --- --- PBC --- --- ---
Authoring in Wikipedia .85 .73 BE1 .83 12.91 .31
BE2 .88 14.32 .22
Note: χ2=229.27, df=152, χ2/df=1.51, RMSEA=.055, NFI=.91, NNFI=.94, CFI=.96, GFI=.88, AGFI=.84.
Table 4 Test of Discriminant Validity
Test # Construct Constrained
model χ2(df) Unconstrained
model χ2(df) Difference Δχ2
Altruism with
1 Expected Reputation 331.37(20) 29.02(19) 302.35(1)
2 Expected Money Reward 320.48(6) 3.06(5) 317.42(1)
3 Attitude 331.31(20) 34.74(19) 296.57(1)
4 Perceived Behavioral Control 315.44(6) 3.07(5) 312.37(1)
5 VC Identity 346.62(20) 28.49(19) 318.13(1)
6 Authoring in Wikipedia 99.38(10) 5.37(9) 94.01(1)
Expected Reputation with
7 Expected Money Reward 368.39(6) 14.92(5) 242.47(1)
8 Attitude 546.85(20) 40.07(19) 506.78(1)
9 Perceived Behavioral Control 415.48(6) 13.41(5) 402.07(1)
10 VC Identity 347.16(20) 54.01(19) 293.15(1)
11 Authoring in Wikipedia 109.47(10) 14.81(9) 94.66(1)
Expected Money Reward with
12 Attitude 507.98(6) 4.93(5) 503.05(1)
13 Perceived Behavioral Control --- --- ---
14 VC Identity 282.10(6) 9.68(5) 272.42(1)
15 Authoring in Wikipedia 91.44(2) 0.58(1) 90.86(1)
Attitude with
16 Perceived Behavioral Control 483.40(6) 4.92(5) 478.48(1)
17 VC Identity 222.34(20) 31.27(19) 191.07(1)
18 Authoring in Wikipedia 98.00(10) 5.93(9) 92.07(1)
Perceived Behavioral Control
19 VC Identity 289.07(6) 10.73(5) 278.34(1)
20 Authoring in Wikipedia 90.93(2) 1.77(1) 89.16(1)
VC Identity with
21 Authoring in Wikipedia 124.20(10) 28.46(9) 95.74(1)
Table 5 Correlation Matrix
Variables Mean Std. 1 2 3 4 5 6 7
1 Altruism 5.05 .94 .73
2 Expected Reputation 4.82 1.15 .21 .80
3 Expected Money Reward 3.21 1.68 .10 .37 ---
4 Attitude 5.89 .80 .27 .31 -.13 .84
5 Perceived Behavioral Control 2.67 .87 .16 .01 -.01 .32 ---
6 VC Identity 4.39 .58 .10 .19 -.32 .58 .28 .74
7 Authoring 1.16 1.12 -.05 -.09 -.23 .16 .45 .27 .85
5.2 Test of Structural Model and Hypotheses
The results of structural model are demonstrated in Figure 2 that all indices show a good fit between model and data. All paths coefficients and t values are also reported in Figure 2 that H1, H3, H4, H6 and H7 are supported, whereas H2 and H5 are not. Additionally, H8 and H9 were tested by means of HRA, hierarchical regression analysis. In the procedure of HRA, attitude and identity, perceived behavioral control, and the interaction items entered the model sequentially, mapping to model 1, 2, and 3. The overall model fit, path coefficients, and difference of R square of each model and its significance are reported in Table 5. The results show that H9 is supported whereas H8 is not. To have the insight of the interaction role-perceived behavioral control (PBC), a plot is exhibited in Figure 3 showing that high PBC has a stronger effect of VC identity on authoring in Wikipedia than low PBC.
Table 6 Results of Hierarchical Regression Analysis for Testing H8 and H9
Variables Model 1 Model 2 Model 3 Note
Attitude -.01(-.11) -.05(-.58) -.02(-.21)
VC Identity .31(3.66) .28(3.23) .25(2.87)
Perceived Behavioral Control .19(2.45) .15(1.88)
PBCAttitude -.06(-.68) H8 is not supported
PBCVC Identity .22(2.30) H9 is supported
ΔR2 .09 .04 .03
F change 8.61*** 6.23** 3.00*
R2 .09 .13 .16
Overall F-Value 8.61*** 8.00*** 6.11***
6. Conclusion
Based on the results of the current research, we find out the relationship between altruism, expected reputation, and expected money reward and both attitude and identity. We also can say that virtual community identity is a more important predictor of authoring than attitude is. It is not sufficient to take action when people think “authoring in Wikipedia is good,” but they praise themselves as a part of the collective, i.e. a member of Wikipedia. Identity to the community is the just force to instigate authoring in Wikipeia.This article explores the motivations determining the volitional authoring of the Wikipedian. Altruism and expected future reputation affect attitude and identity positively, but the relationship between expected money reward and its consequences is negative. We explain the latter that the Wikipedian realize the fact of no money return while they devote themselves to volitional authoring. We also find that the virtual community identity of Wikipedia is more significant to predict volitional authoring than attitude. Besides, the perceived behavioral control, an estimative competence on authoring by author himself, is an accelerator for motivating authoring. Finally, virtual community identity may foster more volitional authoring when the Wikipedian are of higher perception of perceived behavioral control than lower.
Reference
Ajzen, I. (1991), “The theory of planned behavior”, Organizational Behavior and Human Decision Processes, 50(2): 179-211.
Badke, W. (2008), “What to Do with Wikipedia”, http://www.infotoday.com/online/mar08/Badke.shtml.
Bandura, A. (1977), “Self-efficacy: Toward a unifying theory of behavioral change”, Psychological Review, 84(2): 191-215.
Bagozzi, R.P. and Dholakia, U.M. (2002), “Intentional social Action in virtual communities”, Journal of Interactive Marketing, 16(2): 2-21.
Bassellier, G., Benbasat, I. and Reich, B.H. (2003), “The influence of business managers IT competence on championing IT”, Information System Research, 14(4): 317-336.
Bock, G.W., Zmud, R.W., Kim, Y.G. and Lee, J.N. (2005), “Behavioral intention formation in knowledge sharing: Examining the roles of extrinsic motivators”, MIS Quarterly, 29(1): 87-111.
Bock, G.W. and Kim, Y.G. (2002), “Breaking the Myths of Rewards: An Exploratory Study of Attitudes about Knowledge Sharing”, Information Resource Management Journal, 15(2): 14-21.
Chatman, J.A. and Spataro, S.E. (2005), “Using self-categorization theory to understand relational demography-based variations in people's responsiveness to organizational culture”, Academy of Management Journal, 48(2): 321-331.
Chattopadhyay, P. (1999), “Beyond direct and symmetrical effects: The influence of demographic dissimilarity on organizational citizenship behavior”, Academy of Management Journal, 42(3): 273-287.
Chiu, C.M., Hsu, M.H., and Wang, E.T.G. (2006), “Understanding knowledge sharing in virtual communities: An integration of social capital and social cognitive theories”, Decision Support System, 42(3): 1872-1888.
Ciffolilli, A. (2003), “Phantom authority, self-selective recruitment and retention of members in virtual communities: The case of Wikipedia”, FirstMonday, 8(12),
http://firstmonday.org/htbin/cgiwrap/bin/ojs/index.php/fm/article/view/1108
Constant, D.K. and Sproull, L. (1994), “What’s mine is ours, or is it? A study of attitudes about information sharing”, Information Systems Research, 5(4): 400-421.
Donath, J. (1999), “Identity and deception in the virtual community”, In Communities in Cyberspace, Routledge.
Ellemers, N., Kortekaas, P. and Ouwerkerk, J.W. (1999), “Self-categorization, commitment to the group and group self-esteem as related but distinct aspects of social identity”, European Journal of Social Psychology, 29(2-3): 371-389.
Hars, A. and Ou, S. (2002), “Working for free? Motivations for participating in open-source projects”, International Journal of Electronic Commerce, 6(3): 25-39.
Hogg, M.A. and Abrams, D. (1988), Social identifications-A social Psychology of Intergroup Relations and Group Processes, Routledge.
Hsu, C.L. and Chiu, C.M. (2004), “Predicting electronic service continuance with a decomposed theory of planned behaviour”, Behaviour and Information Technology, 23(5): 359-373.
Hsu, C.L. and Lin, J.C.C. (2008), “Acceptance of blog usage: The roles of technology acceptance, social influence and knowledge sharing motivation”, Information and Management, 45(1): 65-74.
Jian, G. and Jeffres, L.W. (2006), “Understanding employees' willingness to contribute to shared electronic databases”, Communication Research, 33(4): 242-261.
Kankanhalli, A., Tan, B.C.Y. and Wei, K.K. (2005), “Contributing knowledge to electronic knowledge repositories: An empirical investigation”, MIS Quarterly, 29(1): 113-143.
Kim, J. (2009), “I want to be different from others in cyberspace-The role of visual similarity in virtual group identity”, Computers in Human Behavior, 25(1): 88-95.
Kollock, P. (1999), “The economies of online cooperation: Gifts and public goods in cyberspace”, in Communities in Cyberspace, Routledge.
Kuo, F.Y. and Young, M.L. (2008), “A study of the intention-action gap in knowledge sharing practices”, Journal of the American Society for Information Science and Technology, 59(8): 1224-1237.
Lerner, J. and Tirole, J. (2002), “Some simple economics of open source”, Journal of Industrial Economics, 50(2): 197-234.
Ma, M. and Agarwal, R. (2007), “Through a glass darkly: Information technology design, identity verification, and knowledge contribution in online communities”, Information Systems Research, 18(1): 42-67.
Nahapiet, J. and Ghoshal, S. (1998), “Social capital, intellectual capital, and the organizational advantage”, Academy of Management Review, 23(2): 242–266.
Nov, O. (2007), “What motivates Wikipedians?” Communications of the ACM, 50(11): 60-64.
Ryan, R.M. and Deci, E.L. (2000), “Intrinsic and extrinsic motivations: Classic definitions and new directions”, Contemporary Educational Psychology, 25(1): 54-67.
Van Dick, R., Grojean, M.W., Christ, O. and Wieseke, J. (2006), “Identity and the extra mile: Relationships between organizational identification and organizational citizenship behaviour”, British Journal of Management, 17(4): 283-301.
Venkatraman, N. (1989), “Strategic orientation of business enterprises: The construct, dimensionality, and measurement”, Management Science, 35(8): 942-962.
Wasko, M.M. and Faraj, S. (2005), “Why should I share? Examining social capital and knowledge contribution in electronic networks of practice”, MIS Quarterly, 29(1): 35-57.
JOSEPH C. SHIH
Department of Information Management, Lunghwa University of Science and Technology, Taoyuan County, Taiwan, (R.O.C.)
Email: joseph@mail.lhu.edu.tw
C. K. FARN
Department of Information Management, National Central University,
Taoyuan County, Taiwan, (R.O.C.)
Email: ckfarn@mgt.ncu.edu.tw
Wikipedia is an online free encyclopedia which is edit by million people spontaneously. This article aims at how Wikipedians’ intrinsic and extrinsic motivations influence their volitional authoring. Research model posits that altruism, expected reputation, and expected money reward affect authoring behavior. More specific, this relationship is also mediated by both attitude and identity. We also regard perceived behavioral control as a critical role for fostering volitional authoring. Sample data were from “Wikipedian Discussion Board” of a famous BBS in Taiwan. All respondents (156 samples) had posted articles in the BBS, but not all had experienced authoring in Wikipedia. Structural equation modeling was used to test the research model. According to the result, we have insight into the motivation of Wikpedian’s volitional authoring.
1. Introduction
Wikipedia, a free online encyclopedia that anyone can edit, has a tremendous impact on how a great many writers gather information about the world (http://www.widipedia.org/). With no paid editors and written by numerous volunteers, Wikipedia is now emerged as the No. 1 go-to information source in the world. Wikipedia also now ranks eighth (July 2009) on the list of most visited sites on the Internet (http://alexa.com/), containing over 2.9 million articles in the English version (July 2009).
However, there are still a few people detracting the value of content (Badke, 2008), but the Wikipedian have persisted in pouring more and more items into the online encyclopedia. The phenomenon is interesting that there are still thousands of people participating in authoring items spontaneously, disregarding the controversial open source project being discredited by non-supporters. Furthermore, the economic exchange perspective posits that an individual’s decision making was found upon rational rule, as benefit surpassing cost. It looks as if the Wikipedian are not consistent with this tenet obviously. This paper, accordingly, is to understand how the Wikipedian’s motivation links to their volitional authoring, more specific, to examine their authoring behavior through the lens of attitude and virtual community identity.
2. Conceptual Background
2.1 Motivations for Participating in Open Source Project
Motivations are commonly categorized into extrinsic and intrinsic by researchers. Intrinsic motivations contain inherent satisfactions rather than their substantive consequence, such as volunteering and enjoying helping others which are congruent with one’s value system, (Kankanhalli et al., 2005; Ryan and Deci, 2000; Wasko and Faraj, 2005). On the other hand, extrinsic motivations means a focus on expected benefits of donating, where the extrinsic rewards are believed to exceed the contribution’s costs (Kankanhalli et al., 2005; Lerner and Tirole, 2002). Motivations of volunteering behavior have noticed in previous research (Clary et al., 1998), for example, Clary et al. (1998) classified six motivational categories for the volunteer. The categories include values, social, understanding, career, protective, and enhancement, and, contrasting with the functions, we describe the Wikipedian’s conceivable motivation in Table 1.
Table 1 Descriptions of Volunteer’s Motivations
Function Conceptual definition Description for Wikipedian
Values The individual volunteers in order to express or act on important values like humanitarianism. The Wikipedian feel it is important to help other by means of authoring.
Understanding The volunteer is seeking to learn more about the world or exercise skills that are often unused. Authoring lets the Wikipedian learn through direct and hands-on activities.
Enhancement One can grow and develop psychologically through volunteer activities. Authoring makes the Wikipedian feel better about themselves.
Career The volunteer has the goal of gaining career-related experience through volunteering. Authoring can help the Wikipedian to get experience related to current job.
Social Volunteering allows an individual to strengthen his or her social relationships. The Wikipedian identify to Wikipedia community.
Protective The individual uses volunteering to reduce negative feelings, such as guilt, or to address personal problems. Authoring is a good escape from day-to-day worry.
2.2 Attitude and Behavioral Control and Online Authoring
Theory of Planned Behavior (TPB) hypothesized that intention to perform a behavior is based on: attitudes, subjective norms and perceived behavioral control (Ajzen, 1991). Attitude towards performing the behavior is defined as a person’s general feeling of performing that behavior if a favorable or unfavorable action. Perceived behavioral control is assumed to reflect past experience as well as anticipated obstacles. The more opportunities and resource that individuals think they possess and the fewer obstacles they anticipate, the greater their perceived control over the behavior. Participating in online activities may be determined by attitude toward and perceived controllability over target behavior, such as blogging, using instant message software, or sharing knowledge in virtual community (Hsu and Chiu, 2008; Kuo and Young, 2008).
2.3 Social Identity and Online Prosocial Behavior
Social identity (SI) captures the main aspects of the individual’s identification with the group in the sense that a participant comes to view himself or herself as “belonging” to a certain group. A person professes “belonging” to a specific group is a psychological state, distinct from being a unique and distinct individual, conferring a collective representation of who one is (Hogg and Abrams, 1988). Many of previous studies related to online volitional behavior have emphasized the significance of virtual community identity. That participants consider themselves as parts of the target online group may strengthen the self-defining relation to the virtual community as well as foster prosocial behavior such like donating knowledge (Chiu et al., 2006; Jian and Jeffres, 2006; Ma and Agarwal, 2007).
3. Research Model and Hypotheses
3.1 Research Model
The present study posits intrinsic and extrinsic motivation associate with both attitude and virtual community identity (VC identity); and attitude, VC identity, and perceived behavior control (PBC) jointly affect authoring behavior. Beside, we also posits that PBC moderates both relationship of attitude-authoring and VC identity-authoring.
3.2 Research Hypotheses
Expressing or acting on important values like humanitarianism (Clary et al., 1998), it may be the most critical motivation of the Wikipedian. People engage in contribution of open content for their own sakes rather than for some external consequences (Hars and Ou, 2002). Positive relation between intrinsic motivation and authoring in Wikipedia was reported in previous study (Nov and Kuk, 2008). Altruism is a variant of intrinsic motivation in which a Wikipedia seeks to increase the welfare of others. Furthermore, altruism have widely held to be associated with positive attitude toward on-line helping behaviors, such as authoring in blog and participating in open source project (Hars and Ou, 2002; Hsu and Lin, 2008). For example, value of altruism, to which the volunteer may reflect their willingness to help people, is the primary function of the Wikipedian’s motivation (Clary and Snyder, 1999). Although Nov and Kuk (2008) proposed general intrinsic motivation in their study, more specific, we argue that altruism influences one’s attitude toward authoring in Wikipedia.
H1: Altruism positively associates with attitude toward authoring.
According to social categorization, people use demographic differences or distinguishable characteristic to categorize one another (Chatman and Spataro, 2005). Wikipedians, obviously different from non-participator, are willing to self-identify themselves as a specific collective, so that those in-group members emerge virtual community identity. We have two reasons to infer the relationship. One reason is that the value system of those who are higher extent of altruism may be more possible to align with the mission of Wikipedia. Altruistic people conceive the principle of unselfish concern for or devotion to the welfare of others, whereas writing items in Wikipedia premises on participant’s volitional behavior. This logic was also argued by Van Dick et al. (2006) that social identification is positively related to prosocial behavior. Another reason is that professional identity is an individual’s self-definition as a member of profession collective (Chreim et al., 2007). They emphasized “role identity” which can make that a professional conducts his or her work look more like professional. Therefore, those high altruistic people are happy to profess themselves as members of Wikipedian. Depending on the virtual community identity, not only a Wikipedian’s professional identity is acknowledged but also his or her altruistic authoring is encouraged. Thus, we propose the following hypothesis.
H2: Altruism positively associates with VC identity.
As economic considerations still critical to the volunteer, helper’s behaviors in cyberspace are as well, Donath (1999) remarked altruism is insufficient to explain helper’s motivation. This viewpoint is even more assertive by Kollock (1999) observing that motivation to contribute to online communities could spring from a variety of sources but none of them depended on altruism. In terms of their viewpoints, the volitional behaviors of the Wikipedian may be influenced by some economic considerations.
Earning reputation is regarded as an extrinsic motivation for a Wikipedian. Reputation is still an important asset, not only in real world but also in cyberspace (Wasko and Faraj, 2005), whereby an individual can leverage to achieve and maintain status within a collective. For the rank of Wikipedian growing with the number of articles edited and accumulated fame as a source of authority, cyberspatial reputation can be a strong incentive to engage authoring activities (Ciffolilli, 2003).
Next, we hypothesize the relationship between money reward and authoring behavior. According to economic exchange theory, one will make decisions by rational self-interest. Thus, prosocial behaviors such as knowledge sharing will occur if its rewards exceed its costs (Bock and Kim, 2002). In light of the rationale, the Wikipedian should have expected money reward while contributing knowledge to Wikipedia. Obviously, the Wikipedian did not conform to the rule. That the Wikipedian plainly understand no money reward but continually author leads to hypothesize the negative relation between money reward and attitude toward authoring. However, it doesn’t mean that Wikipedians hate money but merely they do not expect this extrinsic reward-in terms of money return-while authoring in Wikipedia.
H3: Extrinsic rewards associate with attitude toward authoring
H3a: Expected reputation positively associates with attitude toward authoring.
H3b: Expected money rewards negatively associates with attitude toward authoring.
For earning future reputation in virtual community, a Wikipedian regards social interaction with other aficionados as critical as authoring in virtual community. In additional to engage in authoring and contributing knowledge, the Wikipedian have to maintain a positive social relationship with other Wikipedians. The rationales are that a Wikipedian identifies to his or her community may satisfy the member’s need of self-defining and manifest the status of referent power (Bagozzi and Dholakia, 2002). Earning reputation not only is an antecedent for virtual community identity but also a driven force for authoring. Again, no money return is supposed to authoring in Wikipedia that obtaining identity from the community may, so that we hypothesize a negative relationship between expected money rewards and virtual community identity.
H4: Extrinsic rewards associate with VC identity.
H4a: Expect reputation positively associates with VC identity.
H4b: Expected money rewards negatively associates with VC identity.
Spending time and effort to complete articles, the Wikipedian consider authoring is consistent with their values and positive for online readers. Theory of Reasoned Action posits that behavioral formation is determined by attitude toward the behavior and subject norms of that behavior. The Wikipedian are inducing the authoring behavior while they have a positive attitude toward authoring. Previous studies related to online knowledge sharing have shown the consistent argument (Bock et al., 2005). For example, Kuo and Young (2008) found that the more favorable the individual’s attitude toward knowledge sharing practices, the stronger his/her intention to share knowledge in a teacher’s virtual community.
H5: Attitude toward authoring positively associates with authoring in Wikipedia.
Social identity model of deindividuation effects (SIDE) can demonstrate the relationship between Wikipedian’s community identity and authoring behavior. Volunteering allows an individual to strengthen his or her social relationships (Clay et al., 1998). In light of the rationale, community identity is useful in explaining individuals’ willingness to maintain committed relationship with the group (Nahapiet and Ghoshal, 1998), such as through sharing knowledge within the virtual community (Chiu et al., 2006). Empirical studies have reported that sense of belonging is important and has been used as a test for the presence of an online community. Jian and Jeffres (2006) argues that people are motivated to contribute to shared electronic databases because by doing so they will maintain and affirm relevant identities. Hsu and Lin (2008) argued the influence of social identification for blog users needs the perception of belonging to the virtual community. Likewise, the Wikipedian will present themselves as belonging to the community through continual authoring. Once being in Wikipedian community that suppresses individual differences and emphasizes a common Wikipedian membership, individuals have a high level of group identity and act according to the objective set up by the group (Kim, 2009). Herein, the objective set by Wikipedian community is contributing knowledge, i.e. authoring in Wikipedia.
H6: VC identity positively associates with authoring in Wikipedia.
As a Wikipedian believes himself or herself having sufficient resources to author in the community, the Wikipedian can complete the authoring behavior. Some obstacles such as no available time or incapable of computer skill make authoring impossible. Perceived behavioral control is an important antecedent of behavior theoretically (Ajzen, 1991) that the belief is a form of self-evaluation which influences decisions about what behaviors to undertake (Bandura, 1977). Previous empirical studies of online behavior also have confirmed the notion (Hsu and Chiu, 2004; Kuo and Young, 2008). Accordingly, we propose the following hypothesis.
H7: Perceived behavioral control positively associates with authoring in Wikipedia.
When the Wikipedian are highly perceived control over getting through with new items for Wikipedia, the positive attitude may foster the Wikipedian more authoring behavior. In term of contrast viewpoint-low perceived behavioral control, if a person thinks it good to author in Wikipedia but he does not have time to do this or he does not have enough computer skill to fill out the job, the person will decline the extent of attitude and then impede the authoring behavior. Likewise, when high perceived behavioral control, people with high VC identity will more possibly engaged in authoring in Wikipedia. Thus, we propose the hypothesis of moderating effect.
H8: The relationship between attitude and authoring in Wikipedia is moderated by perceived behavioral control (PBC). More specific, the relationship between attitude and authoring in Wikipedia is stronger when PBC is high than low.
H9: The relationship between VC identity and authoring in Wikipedia is moderated by perceived behavioral control (PBC). More specific, the relationship between VC identity and authoring in Wikipedia is stronger when PBC is high than low.
4. Research Methodology
A questionnaire was deployed on a web site in which allows respondents to answer the questionnaire. Structural equation modeling (SEM) technique was employed to test research model as well as H1 to H7; H8 and H9, the moderating effect of perceived behavioral control on attitude and identity to authoring behavior, were tested by hierarchical regression analysis. Followings are details of the current research methodology.
4.1 Scale Development
The survey questionnaire was designed on the basis of a comprehensive literature review and was refined via several runs of pretests and revisions.
Altruism was measured by a seven-point scale adopted from Chattopadhyay (1999), developed to capture a respondent’s seeking to increase the welfare of others. A sample item is “I will help co-worker (or classmate) who overloads with job (or school-homework).” Expected reputation was measure by a seven-point scale adopted from Constant et al. (1994). Attitude was measured by a seven-point scale adopted from Bock et al. (2005). Virtual community identity was measured by a seven-point scale adopted from Ellemers et al. (1999), developed to capture a respondent’s identification with the virtual community in the sense that the one comes to view himself or herself as a member of Wikipedia community. Perceived behavioral control was measured by a seven-point scale adopted from Armitage et al. (1999), developed to capture a respondent’s ease or difficulty of authoring in Wikipedia. After examining the nature of this scale, we regarded PBC as a formative construct that we would aggregate the score at consequent stage. In order to measure the extent in which the authors contribute to Wikipedia, we operationalized authoring in Wikipedia with two items, one is time consuming per week and another is the frequency of authoring per week. Detail items are omitted due to the limit of space in conference version.
4.2 Data Collection
A web-based site was deployed that respondents could visit to answer the questionnaire which was designed for collecting empirical data of the current study. To have a broad representation of both Wikipedian-authors and non-Wikipedian authors, participants were invited from the Wikipedian’s discussion board of PTT forum which is a famous bulletin board system in Taiwan. We invited them to participate in the survey via an email, in which attaching the web-site’s hyperlink, so that they could visit our web page to answer survey questions. They were also informed that we would donate 5 dollars to Wikipedian Foundation while a questionnaire was finished.
There were totally 181 respondents participating in the survey that the valid responding rate is 86% (156 valid) due to dropping 25 invalid questionnaires. Of 156 samples, the characteristics are demonstrated in Table 2
Table 2 Characteristics of Sample
Gender Education
Male 123 79% Under high school 13 8%
Female 33 21% High school 25 16%
University 80 51%
Authoring experience Graduated school 37 24%
Yes 107 69% PhD 1 1%
No 49 31%
Age
Years of using Wikipedia Under 15 7 5%
Under 1 37 24% 15 to 19 14 9%
1 73 47% 20 to 24 77 49%
2 32 21% 25 to 29 40 26%
3 9 6% 30 to 34 13 8%
4 4 3% 35 to 40 3 2%
5 and above 1 1% 40 above 2 1%
5. Data Analysis
5.1 Test of Measurement Model
Initial results of the CFA indicates that model were not fit the data well. A careful and iterative inspection of LISREL output revealed that some items did not load on the designated latent factors appropriately, such as standardized loading < 0.6 or associated with high modification indices. All indices are above cut-off value except GFI is slightly lower. The detail results are demonstrated in Table 3.
Before testing the structural model, it is also necessary to examine whether the measurement model had a satisfactory level of validity and reliability. While CR is greater than 0.7 and AVE is greater than 0.5, it implies that the variance captured by the latent construct is more than that by error component (Bagozzi et al., 1991). That is, each measure is accounting for 50 percent or more of the variance of the underlying latent variable (Chin, 1998). As the reports in Table 3, CRs and AVEs are all above recommended cut-off values that the scale is of internal consistency reliability.
Convergent validity ensures that all items measure a single latent construct, and it is established if all item loadings are greater than or equal to the recommended cut-off level of 0.70 (Bassellier et al., 2003). Our results showed that almost loadings of each latent variable are above the cut-off value that only two items’ are slightly lower. The details are also exhibited in Table 3.
Discriminant validity reflects the level to which the measures for each dimension are distinctively different from each other. We applied the chi-square difference test to assess the discriminant validity of the measurement model (Bassellier et al., 2003). Accordingly, we conducted 15 pair-wise tests (six constructs) that the results are reported in Table 4. All Δχ2 differences are significant above the level of Pr[χ2(1) ≥ 3.84]=0.05, indicating strong support for discriminant validity (Bassellier et al., 2003; Venkatraman, 1989). Additionally, the correlation matrixes are reported in Table 5.
Table 3 Results of Measurement Model Test
Factor Composition Reliability AVE Items Loading t-value Error
Altruism .84 .57 AL1 .63 8.67 .60
AL2 .71 10.17 .49
AL3 .92 14.28 .16
AL4 .73 10.43 .47
Expected Reputation .88 .64 ER1 .76 11.26 .43
ER2 .91 14.86 .17
ER3 .88 14.01 .23
ER4 .63 8.92 .60
Expected Money Reward --- --- EMR --- --- ---
Attitude .91 .71 AT1 .78 11.92 .39
AT2 .83 12.89 .32
AT3 .89 14.39 .22
AT4 .87 13.87 .25
Virtual Community Identity .83 .55 CI1 .72 10.09 .49
CI2 .74 10.58 .45
CI3 .74 10.60 .45
CI4 .76 11.00 .42
Perceived Behavioral Control --- --- PBC --- --- ---
Authoring in Wikipedia .85 .73 BE1 .83 12.91 .31
BE2 .88 14.32 .22
Note: χ2=229.27, df=152, χ2/df=1.51, RMSEA=.055, NFI=.91, NNFI=.94, CFI=.96, GFI=.88, AGFI=.84.
Table 4 Test of Discriminant Validity
Test # Construct Constrained
model χ2(df) Unconstrained
model χ2(df) Difference Δχ2
Altruism with
1 Expected Reputation 331.37(20) 29.02(19) 302.35(1)
2 Expected Money Reward 320.48(6) 3.06(5) 317.42(1)
3 Attitude 331.31(20) 34.74(19) 296.57(1)
4 Perceived Behavioral Control 315.44(6) 3.07(5) 312.37(1)
5 VC Identity 346.62(20) 28.49(19) 318.13(1)
6 Authoring in Wikipedia 99.38(10) 5.37(9) 94.01(1)
Expected Reputation with
7 Expected Money Reward 368.39(6) 14.92(5) 242.47(1)
8 Attitude 546.85(20) 40.07(19) 506.78(1)
9 Perceived Behavioral Control 415.48(6) 13.41(5) 402.07(1)
10 VC Identity 347.16(20) 54.01(19) 293.15(1)
11 Authoring in Wikipedia 109.47(10) 14.81(9) 94.66(1)
Expected Money Reward with
12 Attitude 507.98(6) 4.93(5) 503.05(1)
13 Perceived Behavioral Control --- --- ---
14 VC Identity 282.10(6) 9.68(5) 272.42(1)
15 Authoring in Wikipedia 91.44(2) 0.58(1) 90.86(1)
Attitude with
16 Perceived Behavioral Control 483.40(6) 4.92(5) 478.48(1)
17 VC Identity 222.34(20) 31.27(19) 191.07(1)
18 Authoring in Wikipedia 98.00(10) 5.93(9) 92.07(1)
Perceived Behavioral Control
19 VC Identity 289.07(6) 10.73(5) 278.34(1)
20 Authoring in Wikipedia 90.93(2) 1.77(1) 89.16(1)
VC Identity with
21 Authoring in Wikipedia 124.20(10) 28.46(9) 95.74(1)
Table 5 Correlation Matrix
Variables Mean Std. 1 2 3 4 5 6 7
1 Altruism 5.05 .94 .73
2 Expected Reputation 4.82 1.15 .21 .80
3 Expected Money Reward 3.21 1.68 .10 .37 ---
4 Attitude 5.89 .80 .27 .31 -.13 .84
5 Perceived Behavioral Control 2.67 .87 .16 .01 -.01 .32 ---
6 VC Identity 4.39 .58 .10 .19 -.32 .58 .28 .74
7 Authoring 1.16 1.12 -.05 -.09 -.23 .16 .45 .27 .85
5.2 Test of Structural Model and Hypotheses
The results of structural model are demonstrated in Figure 2 that all indices show a good fit between model and data. All paths coefficients and t values are also reported in Figure 2 that H1, H3, H4, H6 and H7 are supported, whereas H2 and H5 are not. Additionally, H8 and H9 were tested by means of HRA, hierarchical regression analysis. In the procedure of HRA, attitude and identity, perceived behavioral control, and the interaction items entered the model sequentially, mapping to model 1, 2, and 3. The overall model fit, path coefficients, and difference of R square of each model and its significance are reported in Table 5. The results show that H9 is supported whereas H8 is not. To have the insight of the interaction role-perceived behavioral control (PBC), a plot is exhibited in Figure 3 showing that high PBC has a stronger effect of VC identity on authoring in Wikipedia than low PBC.
Table 6 Results of Hierarchical Regression Analysis for Testing H8 and H9
Variables Model 1 Model 2 Model 3 Note
Attitude -.01(-.11) -.05(-.58) -.02(-.21)
VC Identity .31(3.66) .28(3.23) .25(2.87)
Perceived Behavioral Control .19(2.45) .15(1.88)
PBCAttitude -.06(-.68) H8 is not supported
PBCVC Identity .22(2.30) H9 is supported
ΔR2 .09 .04 .03
F change 8.61*** 6.23** 3.00*
R2 .09 .13 .16
Overall F-Value 8.61*** 8.00*** 6.11***
6. Conclusion
Based on the results of the current research, we find out the relationship between altruism, expected reputation, and expected money reward and both attitude and identity. We also can say that virtual community identity is a more important predictor of authoring than attitude is. It is not sufficient to take action when people think “authoring in Wikipedia is good,” but they praise themselves as a part of the collective, i.e. a member of Wikipedia. Identity to the community is the just force to instigate authoring in Wikipeia.This article explores the motivations determining the volitional authoring of the Wikipedian. Altruism and expected future reputation affect attitude and identity positively, but the relationship between expected money reward and its consequences is negative. We explain the latter that the Wikipedian realize the fact of no money return while they devote themselves to volitional authoring. We also find that the virtual community identity of Wikipedia is more significant to predict volitional authoring than attitude. Besides, the perceived behavioral control, an estimative competence on authoring by author himself, is an accelerator for motivating authoring. Finally, virtual community identity may foster more volitional authoring when the Wikipedian are of higher perception of perceived behavioral control than lower.
Reference
Ajzen, I. (1991), “The theory of planned behavior”, Organizational Behavior and Human Decision Processes, 50(2): 179-211.
Badke, W. (2008), “What to Do with Wikipedia”, http://www.infotoday.com/online/mar08/Badke.shtml.
Bandura, A. (1977), “Self-efficacy: Toward a unifying theory of behavioral change”, Psychological Review, 84(2): 191-215.
Bagozzi, R.P. and Dholakia, U.M. (2002), “Intentional social Action in virtual communities”, Journal of Interactive Marketing, 16(2): 2-21.
Bassellier, G., Benbasat, I. and Reich, B.H. (2003), “The influence of business managers IT competence on championing IT”, Information System Research, 14(4): 317-336.
Bock, G.W., Zmud, R.W., Kim, Y.G. and Lee, J.N. (2005), “Behavioral intention formation in knowledge sharing: Examining the roles of extrinsic motivators”, MIS Quarterly, 29(1): 87-111.
Bock, G.W. and Kim, Y.G. (2002), “Breaking the Myths of Rewards: An Exploratory Study of Attitudes about Knowledge Sharing”, Information Resource Management Journal, 15(2): 14-21.
Chatman, J.A. and Spataro, S.E. (2005), “Using self-categorization theory to understand relational demography-based variations in people's responsiveness to organizational culture”, Academy of Management Journal, 48(2): 321-331.
Chattopadhyay, P. (1999), “Beyond direct and symmetrical effects: The influence of demographic dissimilarity on organizational citizenship behavior”, Academy of Management Journal, 42(3): 273-287.
Chiu, C.M., Hsu, M.H., and Wang, E.T.G. (2006), “Understanding knowledge sharing in virtual communities: An integration of social capital and social cognitive theories”, Decision Support System, 42(3): 1872-1888.
Ciffolilli, A. (2003), “Phantom authority, self-selective recruitment and retention of members in virtual communities: The case of Wikipedia”, FirstMonday, 8(12),
http://firstmonday.org/htbin/cgiwrap/bin/ojs/index.php/fm/article/view/1108
Constant, D.K. and Sproull, L. (1994), “What’s mine is ours, or is it? A study of attitudes about information sharing”, Information Systems Research, 5(4): 400-421.
Donath, J. (1999), “Identity and deception in the virtual community”, In Communities in Cyberspace, Routledge.
Ellemers, N., Kortekaas, P. and Ouwerkerk, J.W. (1999), “Self-categorization, commitment to the group and group self-esteem as related but distinct aspects of social identity”, European Journal of Social Psychology, 29(2-3): 371-389.
Hars, A. and Ou, S. (2002), “Working for free? Motivations for participating in open-source projects”, International Journal of Electronic Commerce, 6(3): 25-39.
Hogg, M.A. and Abrams, D. (1988), Social identifications-A social Psychology of Intergroup Relations and Group Processes, Routledge.
Hsu, C.L. and Chiu, C.M. (2004), “Predicting electronic service continuance with a decomposed theory of planned behaviour”, Behaviour and Information Technology, 23(5): 359-373.
Hsu, C.L. and Lin, J.C.C. (2008), “Acceptance of blog usage: The roles of technology acceptance, social influence and knowledge sharing motivation”, Information and Management, 45(1): 65-74.
Jian, G. and Jeffres, L.W. (2006), “Understanding employees' willingness to contribute to shared electronic databases”, Communication Research, 33(4): 242-261.
Kankanhalli, A., Tan, B.C.Y. and Wei, K.K. (2005), “Contributing knowledge to electronic knowledge repositories: An empirical investigation”, MIS Quarterly, 29(1): 113-143.
Kim, J. (2009), “I want to be different from others in cyberspace-The role of visual similarity in virtual group identity”, Computers in Human Behavior, 25(1): 88-95.
Kollock, P. (1999), “The economies of online cooperation: Gifts and public goods in cyberspace”, in Communities in Cyberspace, Routledge.
Kuo, F.Y. and Young, M.L. (2008), “A study of the intention-action gap in knowledge sharing practices”, Journal of the American Society for Information Science and Technology, 59(8): 1224-1237.
Lerner, J. and Tirole, J. (2002), “Some simple economics of open source”, Journal of Industrial Economics, 50(2): 197-234.
Ma, M. and Agarwal, R. (2007), “Through a glass darkly: Information technology design, identity verification, and knowledge contribution in online communities”, Information Systems Research, 18(1): 42-67.
Nahapiet, J. and Ghoshal, S. (1998), “Social capital, intellectual capital, and the organizational advantage”, Academy of Management Review, 23(2): 242–266.
Nov, O. (2007), “What motivates Wikipedians?” Communications of the ACM, 50(11): 60-64.
Ryan, R.M. and Deci, E.L. (2000), “Intrinsic and extrinsic motivations: Classic definitions and new directions”, Contemporary Educational Psychology, 25(1): 54-67.
Van Dick, R., Grojean, M.W., Christ, O. and Wieseke, J. (2006), “Identity and the extra mile: Relationships between organizational identification and organizational citizenship behaviour”, British Journal of Management, 17(4): 283-301.
Venkatraman, N. (1989), “Strategic orientation of business enterprises: The construct, dimensionality, and measurement”, Management Science, 35(8): 942-962.
Wasko, M.M. and Faraj, S. (2005), “Why should I share? Examining social capital and knowledge contribution in electronic networks of practice”, MIS Quarterly, 29(1): 35-57.
訂閱:
文章 (Atom)