The CEO Project: An Introduction Chris Partridge Technical Report 07/02, LADSEB-CNR Padova, Italy, December 2002 Questo lavoro è stato condotto nell'ambito dell'attività del gruppo di ricerca “Modellazione concettuale e Ingegneria della Conoscenza” del LADSEB-CNR LADSEB-CNR Corso Stati Uniti 4 I-35127 PADOVA (PD) e-mail: mbox@ladseb.pd.cnr.it fax: +39 049-829.5763 tel: +39 049-829.5702 -- 1 of 43 -- -- 2 of 43 -- The CEO Project: An Introduction Page 3 © Chris Partridge 2002. All rights reserved Chris Partridge1 & 2 1 BORO Program, England partridge@BOROProgram.org http://www.BOROProgram.org/ 2 National Research Council, LADSEB-CNR, Italy partridge@ladseb.pd.cnr.it http://www.ladseb.pd.cnr.it/infor/ontology/BusinessObjectsOntology.html Abstract: This is, in essence, the project initiation paper for the CEO Project. Its main concern is explaining the project’s aims, how it intends to achieve them and the methodological framework within which the project will work. It explains the origins, conception and motivation for the project and gives an outline of the management framework for the project, in particular the first synthesis stage. It clarifies the terms of the art and describes the nature of ontological analysis. It also characterises the requirements that shape it and the meta-ontological choices and analytic styles that underlie it. Finally, it describes the potential applications and the next steps. Introduction This is, in essence, the project initiation paper for the CEO (Core Enterprise Ontology) Project. Its primary purpose is to explain the project’s aims and how it intends to achieve them. It also describes the methodological framework within which the project will work. It can also be used to give interested parties an introduction to the project. The first sections start by explaining the origins, conception and motivation for the project. They then give an outline of the management framework of the project, in particular the first synthesis stage. The next sections deal with the approach adopted. This is based upon ontology, a millennia old discipline within philosophy. But its application to enterprise computing is both innovative and radical. Hence it needs some initial explanation. This is given by clarifying the terms of the art and then describing the nature of ontological analysis. The adopted approach is characterised firstly in terms of the requirements that shape it and by the meta-ontological choices and analytic styles that underlie it. The final sections describe the potential applications and the next steps. Origins of the CEO Project The CEO Project has its origins in the REV-ENG Methodology. Origins of the REV-ENG Methodology The REV-ENG methodology grew out of a series of legacy application re-engineering projects each of which started with the re-engineering of a business model out of the existing legacy application1 . What differentiated these projects’ re-engineering approach was a focus on recovering the ontological model of the business objects that underlay the legacy applications. This was typically a demanding task as there was 1 More of the history can be found in the Preface to Partridge (1996) Business Objects: Re - Engineering for re - use. -- 3 of 43 -- The CEO Project: An Introduction Page 4 © Chris Partridge 2002. All rights reserved little or no documentation, only the implemented application. Over time the approach crystallised into a systematic process, which was codified into the REV-ENG (for REVerse ENGineering) Methodology: this is thoroughly documented in (Partridge 1996). Comparisons between the early projects using the REV-ENG ontological analysis revealed that a number of the same general patterns were being repeatedly unearthed – surprisingly even in quite different business areas (e.g. banking and telecommunications). Typically the specific patterns in the applications looked different because they were combinations of different sets of general patterns. It soon became clear that significant time was being wasted repeatedly re-engineering these from scratch. This indicated the potential for high levels of re-use, which was initially exploited by making the general patterns available for re-use in subsequent projects2 . Experience also showed that the potential for generalising (and so simplifying) was rarely exhausted in a single re-engineering. The general patterns found in one project were found, in subsequent projects, to be combinations of even more general patterns. This indicated that there was significant scope for evolving the patterns to greater and greater levels of generality and simplicity. The conception of the CEO Project The CEO Project was conceived out of the realisation that there would be significant economies in the application of the REV-ENG approach if one could start with a core business model for the enterprise. Reasonable economies would come from having the lower level general patterns found in single re-engineerings. But the really significant benefits would come from the very general patterns found in heavy duty re-engineering. Hence the CEO’s overall aim is not only to recover the most common patterns found in businesses into a coherent and consistent model – but also to try and evolve this to much higher levels of generality and simplicity. Motivation for the CEO Project The development of the CEO is motivated by the expectation that building an ontological model with patterns that have high levels of generality and simplicity will significantly enhance the benefits of using the bare REV-ENG methodology: both further reducing the costs and significantly increasing the benefits that accrue at the business requirements (more specifically the business modelling) stage of projects. Costs will come down as projects re-use the CEO’s ready-made foundation instead of building from scratch. Benefits of using the REV-ENG approach will increase in a number of areas, the main ones being: • Reducing complexity, • Improved inter-operability. • Increasing longevity, and • Technology proofing, Reducing complexity: The approach used by the CEO enables increases in the generality of the business model that lead to both significant reductions in complexity and increases in functionality (measured against the legacy system). The key lesson is that complexity is not inherent – it is apparent and much of it can be re-engineered away. The complexity of the business model is a significant contributor to the overall 2 Described in pp.276-8 of Ibid.. -- 4 of 43 -- The CEO Project: An Introduction Page 5 © Chris Partridge 2002. All rights reserved complexity in large business applications – and so development and maintenance costs. The complexity cost of increasing functionality is a major barrier to improvements. Reducing these has paybacks all the way through the lifetime of the application. Improving inter-operability: The CEO model will provide a canonical picture of the business that can be used as an inter-lingua for communicating between applications. Where applications have been re-engineered using the CEO model, their inter- operability will be much simpler 3 . Increasing longevity: Experience to date seems to show that underlying the apparently changing forms of business software there are a relatively stable set of general patterns. The changes in business practices are often merely different combination of these patterns. Building an application based upon these can significantly increase its longevity. From an investment perspective, this increase in the term, leads to a corresponding increase in returns. Technology proofing: An ontology’s focus is on the business rather than the application. The ontology model will represent business objects independently of the technology that is used to implement it. The prime benefit of this independence is an asset that is future proofed against technological innovation. A management framework for the CEO Project To help focus the CEO work, a framework for the project has been established. These three ‘management’ elements of the framework are described below: • the prime goal and deliverable, • the scope, and • an initial project work plan. Setting the CEO’s prime goal and deliverable Given its motivation, the CEO Project’s prime goal is to exploit the ontological approach initially codified in the REV-ENG Methodology to develop a toolkit that enterprises can use to both reduce the costs and significantly increase the benefits of producing a business model. The CEO’s prime deliverable, and the main element of its toolkit, will be an ontological model4 . This will represent the objects that exist in the enterprise field in a standard way. Setting the CEO’s scope The scope of the CEO is circumscribed at these three levels: • ONTOLOGY, • CORE ontology, • core ENTERPRISE ontology. The scope of the initial analysis work is also circumscribed, to an extent, by the scope of the applications included within the CEO’s re-engineering approach. 3 For more details see Partridge (2002b) The Role of Ontology in Integrating Semantically Heterogeneous Databases and Partridge (2002f) What is a customer?. 4 The terms ‘ontology’, ‘ontological model and ‘epistemology’ are described later. -- 5 of 43 -- The CEO Project: An Introduction Page 6 © Chris Partridge 2002. All rights reserved The scope of an ONTOLOGY The scope of the CEO is focused on the application’s ontology – the business objects it refers to independently of the way it represents them. There is a substantial body of philosophical work that provides a framework for talking about business objects in this way. Good starting points are Quine’s notion of ontological commitment5 and Armstrong’s notion of truthmaker 6 . In looking at the way an application represents the business, we can ask what this representation is ontological committed to – what objects is it committed to saying exist? Similarly we can ask what are this representation’s truthmakers - what objects make the application’s representations true? These are the objects that it acknowledges exist (or can exist). An ontological model is a model of these. The CEO’s ontological model is intended to represent the objects that exist in its field – the enterprise. The epistemological aspects of this domain – covering what and how applications may know what they do – are not within scope. The scope of a CORE ontology What is a ‘Core’ Ontology? (Breuker, Valente et al. 1997) provide this useful description – “an intermediary between the generic top and the domain ontologies that contain the categories that define what a field is about.” Where a “field is a discipline, industry or area of practice that unifies many application domains …”. The stratification and segmentation of ontologies into fields and domains (or whatever) is a practical matter – and is guided, as Breuker et al suggest, by how much it helps in practice to provide a unifying structure. The key point is that given a ‘field’ (or domain) there are core ‘categories’ that help to “define what [it] is about.” This the two key characteristics of a core ontology are generality and unity. The focus of the CEO is on the unifying categories for its selected field – the enterprise. The scope of the CEO cannot however be restricted to purely these categories. For the ontological model to make sense, it needs to be embedded in a top ontology and fleshed out with some domain elements. The scope of a core ENTERPRISE ontology The CEO’s field is the enterprise, and its ontological model will focus on the major categories that ‘define what this field is about’. At this initial stage of the project, making a rough intuitive guess at what these might be helps to bring the scope of the CEO into focus and a basis for organising the initial analysis work. Intuitively these three seem like the most suitable major categories: • Person (AKA Party), who can enter into a • Transaction, which often include agreements which involve an • Asset. Experience with the REV-ENG methodology suggests that the recovered ontology is likely to give a radically different perspective on these categories – which could be transformed by the analysis. 5 See Quine (1964) Word and object. 6 See Armstrong (1997b) A world of states of affairs. -- 6 of 43 -- The CEO Project: An Introduction Page 7 © Chris Partridge 2002. All rights reserved Broad categories All three categories are broad. For example the first category, Person, encompasses both people and organisations. One of its unifying characteristics is that its members can enter into transactions. Hence the name in the data modelling community for Persons – Parties – as in ‘a party to a contract’7 . Intertwined categories The unified nature of the enterprise field means that the categories are closely interlinked. This involves a network of ontological dependence. Transactions (contracts) are always entered into by persons – the transaction cannot exist unless the person does. This makes them (ontologically) dependent upon Persons for their existence. Transactions also typically involve and so are dependent upon assets, in a sense in which these are not dependent upon transactions8 . Similarly assets are owned by persons – and so dependent upon them. But the unification is much closer, more intertwining, with the categories overlapping – for example, a company seems to be both a person that can enter into agreements and an asset that can be bought and sold. It is also, from the perspective of incorporation, seems to be a kind of contract (transaction). The completed CEO will properly account for this intertwining: both dependence and overlapping. The scope of the analysis work The scope of the analysis is circumscribed by its re-engineering approach. The contents of the ontologies and applications provide an input to the analysis that helps to set boundaries on the analysis. This helps to ensure that the analysis focuses on content relevant to the kinds of systems that enterprises currently deploy. Provided the sample is reasonably large, this approach does not exclude relevant content that is not in the input sample. The analysis is looking for the general patterns that underlie the existing content. All the relevant patterns are likely to be exhibited in even a small sample, though they may be easier to extract from a reasonably large one. The ontological model will allow for these to be combined in ways that are not exemplified in the sample giving a content that provides a reasonable coverage of the domain – certainly one far exceeding the input sample. Drawing up the initial project work plan At this initial stage a broad project work plan has been drawn up. This envisages two major stages: • A synthesis stage – A Synthesis of (selected) State of the Art Enterprise Ontologies (SSAEO) to produce a Base Enterprise Ontology (BEO), which will act at the foundation for the second stage. • A development stage – A development of the BEO into the industrial strength CEO ontology. 7 As is done, for example, in Inmon, et al. (1997) The data model resource book. 8 Assets are not obviously dependent upon transactions, as the legal notion of inalienable assets, ones that cannot be sold, illustrates. -- 7 of 43 -- The CEO Project: An Introduction Page 8 © Chris Partridge 2002. All rights reserved The synthesis (SSAEO) stage The initial stage has been planned in more detail than subsequent stages. It has the goal of harvesting the insights from the best of breed enterprise ontologies and their synthesis into a single coherent whole. The initial informal review The SSAEO started with an informal review of enterprise ontologies with two goals: • An assessment of the state of the art, and • A selection of the best of breed ontologies to be synthesised. An assessment of the state of the art The informal review found that the ‘state of the art’ is immature, in particular that: • there are not many enterprise ontologies (though there are many resources from which these could be mined), and • those that exist have not yet reached ‘industrial strength’ as ontologies for semantic interoperability. This second point is one of the reasons why the SSAEO synthesis needed to be more of a full re-engineering rather than a straight-forward merge/integration9 . A selection of the best of breed ontologies The review selected the following ontologies for synthesis: • TOronto Virtual Enterprise - TOVE (Fox, Barbuceanu et al. 1996) (Fox, Chionglo et al. 1993) (TOVE:http), • AIAI’s Enterprise Ontology - EO (Uschold, King et al. 1997) (Uschold, King et al. 1998) (EO:http), • Cycorp’s Cyc® Knowledge Base – CYC (Lenat and Guha 1990) (CYC:http), and • W.H. Inmon’s Data Model Resource Book - DMRB 10 (Inmon, Silverston et al. 1997). Task breakdown for the SSAEO analysis The SSAEO involves a substantial amount of work. As such, it made sense to break it down into a smaller number of tasks. It was decided that it should be broken down along two dimensions – the selected ontologies and the core categories. The first area selected for analysis was the TOVE ontology and the Person core category – this task 9 And so does not fall neatly into one of the usual categories; for example, those of integration, merge and use in Gomez-Perez, et al. (1999) Some Issues on Ontology Integration, that assume an underlying homogeneity among the ontologies. 10 In its own terms, this is a universal data model. However, from our perspective, it is in many respects an ontology. We considered having a number of commercial data models in the sample, but found that they were very similar – so there would be no real benefit. Inmon, et al. (1997) The data model resource book and Hay (1996) Data model patterns were neck and neck as the commercial data model representative. We selected Inmon (1997) as it seemed slightly more accessible. Note, a two volume revised edition of this has appeared: Silverston (2001a) The data model resource book 1, Silverston (2001b) The data model resource book 2. -- 8 of 43 -- The CEO Project: An Introduction Page 9 © Chris Partridge 2002. All rights reserved was named Synthesis of a TOVE Person Ontology (STPO). The segmented tasks and the two dimensions are shown diagrammatically in Figure 1 below. EO CYC DMRB CEO TOVE Persons Synthesis Stage Development Stage R E V I E W STPO Overall CEO Project Transactions Assets Figure 1 – The synthesis stage of the CEO project Clarifying the terms of art The key deliverable for the CEO is an ontological model – and this is based upon the notion of an ontology. The CEO needs to clarify what it means by these and related terms as, in the last few decades, the kinds of things that have been called ontologies has increased at least ten-fold. This clarification starts with two basic terms: ontology and semantics. Ontology Central to the CEO’s approach is the traditional philosophical (metaphysical) notion of ontology – where this is “the set of things whose existence is acknowledged by a particular theory or system of thought.” 11 Here the set of things is not just restricted to simple entities, it includes every type of thing that exists: for example, it can include relations and/or states of affairs, if these are deemed to exist. This view was famously summarised by Quine, who claimed that the question ontology asks can be stated in three words ‘What is there?’ – and the answer in one ‘everything’. Not only that, but tongue in cheek, he also said “everyone will accept this answer as true” though he admitted that there was some more work to be done as “there remains room for disagreement over cases.” 12 Quine’s glib description captures the common intuitive position of many systems analysts, who unthinkingly assumed that the answer to the question “What is there – according to this application?” will be the set of things that the application represents. Within the IT community there is no technique for identifying this ‘set of things’ apart from intuition. However, there is substantial body of philosophical work that provides techniques for analysing objects in this way. As noted earlier, good starting points are Quine’s notion of ontological commitment and Armstrong’s notion of truthmaker. In looking at the way a scheme represents its domain, we can ask ‘What is the 11 E. J. Lowe in the Oxford Companion to Philosophy. 12 In Quine (1948) On what there is, reprinted in Quine (1980) From a logical point of view. -- 9 of 43 -- The CEO Project: An Introduction Page 10 © Chris Partridge 2002. All rights reserved ontological commitment of this representation?’ – ‘What objects is it committed to saying exist?’ Similarly, we can ask ‘What things make the representation true?’ In this way, one can clearly differentiate between how something is represented (the representation) and what is being represented (the ontology). These can be (and often are) quite different, and different applications often have quite different representations. Some care needs to be taken to distinguish this traditional metaphysical use of the word ‘ontology’ from one that has recently developed in some parts of Computer Science. Here an ontology is regarded as a “specification of a conceptualisation” (Gruber 1993) and has been applied to a wide range of things, including dictionaries. This sense of the word does not give a fine-grained enough tool for the CEO’s needs. For example, it regards an application as simply an ontology – and so it cannot make sense of talking about the ontology underlying it, let alone underlying a group of applications with a footprint over the same domain. A similar point can be made about conceptual schemas, such as that described in ANSI/X3/SPARC (Tsichritzis and Klug 1978). These deal with representations of the conceptual perspective, and reflect how we conceive of the world – which is, in ways important for business modelling, not quite the same as what our conceptualisation commits to existing in the world (or what things make the conceptualisation true). It has been recognised for a long time that metaphysical ontology has a role to play in IT. Over thirty years ago, (Mealy 1967) suggested that it was essential. (Kent 1978) makes a similar point at book length. However, it was only in the 1990’s that interest started to really grow, particularly in AI. However, work in this area has tended to be done using a revised conceptual notion of ontology – which is not suitable for the CEO’s purposes. In the sample of best of breed ontologies chosen by the CEO, the AI based ones (TOVE, EO and CYC) fall into this category. The DMRB, which (unlike the AI ontologies) is a distillation of actual practice, tries to look through the representation to the objects being represented – and, as such, takes a view reasonably consistent with metaphysical ontologies. Semantics Along with the traditional philosophical sense of ontology there is a related notion of semantics – where this is the relationship between words (data) and the world – the things the words (data) describe13 . This needs to be distinguished from the different, but related, sense of the word in linguistics where it means the study of meaning14 . These notions of ontology and semantics can then be used to describe three other useful notions – that of an ontological model, canonical scheme and semantic divergence. Ontological model Someone who takes the metaphysical view needs to have a way to describe the ontology. At a bare minimum, they can make an inventory of the objects. As this includes relations, the result is more like a model than a mere list – so it is not 13 Or as Nelson Goodman put it in his Introduction to Quine (1973) The roots of reference – “… an important relation of words to objects – or better – of words to other objects, some of which are not words – or even better, of objects some of which are words to objects some of which are not words.” 14 “Semantics – the study of meaning” from the Concise Oxford Dictionary of Linguistics, © Oxford University Press 1997. -- 10 of 43 -- The CEO Project: An Introduction Page 11 © Chris Partridge 2002. All rights reserved stretching the truth to call this an ontological model. For practical reasons, the model cannot name every object – and so typically restricts itself to naming types of objects and a representative sample of instances. What characterises an ontological model is that it directly reflects the ontology. There is a simple semantics where each object in the ontology has a direct relationship with the corresponding representation in the model 15 . One of the characteristics of an ontological model is that the representations in it can be regarded as the names of the objects in the ontology – from a Fregean perspective as reference and no sense (from a Millian perspective as denotation without connotation). In (Marcus 1993), Ruth Barcan Marcus (explicitly following in the footsteps of Mill and Russell16 ) calls this ‘tagging’. The distinction between an ontology and its ontological model should now be clear. However, ‘ontological model’ is a cumbersome term and it is usually clear from the context whether the ontology or its model is being referred to. So from now on, where the context can determine this, the term ontology will be used. Semantic heterogeneity Most applications are not ontological models. This is plain from a phenomenon commonly found in applications and much discussed in database literature – semantic heterogeneity. (Sheth and Larson 1990), on p. 187, provide a description of it. They suggest that heterogeneity occurs “… when there is a disagreement about the meaning, interpretation or intended use of the same or related data [in different databases].” But they note that “… this problem is poorly understood, and there is not even an agreement regarding a clear definition of the problem.” From an ontological perspective it can be described as two semantically different representations of the same objects. Clearly where there is semantic heterogeneity both (all) of the representations cannot be ontological models. Design automomy and diversity Sheth and Larson (among others) note that a prime source of semantic heterogeneity is what they call design automomy. They describe this (on p. 187 of (Sheth and Larson 1990)) as “the ability of a component DBS to choose its own design with respect to any matter”. As they note, this includes “The conceptualization or semantic interpretation of the data (which greatly contributes to the problem of semantic heterogeneity)”. In fact, they say: “Heterogeneity [in general] … is primarily caused by design autonomy among component DBSs.” Of course, autonomy by itself does not lead to heterogeneity. There is in principle no reason why two autonomous designers should not end up with the same design. However, in practice, autonomy allows what I have called design diversity17 to manifest itself – where this is the actual manifestation of two different designs for the same objects. This diversity is partly the result of the different requirements of the 15 This is called strong reference within the REV-ENG Methodology described in Partridge (1996) Business Objects: Re - Engineering for re - use. See also ” Russell and Blackwell (1983) The Collected Papers of Bertrand Russell, Vol 8: p.176: “In a logically perfect language, there will be one word and no more for every simple object”. 16 Mill (1848) A system of logic and Russell (1919) Introduction to mathematical philosophy. 17 See Partridge (2002b) The Role of Ontology in Integrating Semantically Heterogeneous Databases and Partridge (2002e) The Role of Ontology in Semantic Integration. -- 11 of 43 -- The CEO Project: An Introduction Page 12 © Chris Partridge 2002. All rights reserved applications. But it also, partly, the result of the large amount of judgment exercised by the designers. This is reflected in the fact that different designers will typically (as a result of different judgements, different trade-offs) come up with different application designs for two similar applications. It can be quite surprising how different the designs can be 18 . Semantic divergence The notion of semantic heterogeneity is not based upon an ontological perspective. From this perspective the more relevant phenomenon is semantic divergence. This occurs where the semantic relationship between the ontology and the representation is not direct and straightforward. This is related to the notion of ontological model – as these have no semantic divergence. The kind of ontological analysis proposed for developing the CEO involves the extraction of an ontological model from applications, and this can be characterised as identifying and removing semantic divergences. Classic example of semantic divergence Semantic divergence in a common feature of our representations, including applications. A classical example that is often used to illustrate it is data that represents the average family as having 2.4 children. In answer to the question ‘How is this represented?’ – the answer is as a family with children. The answer to the question ‘What is being represented?’ (or what is being ontologically committed to, or what makes the representation true) is quite different. It is not, as the outward form suggests, a family – but a relationship between a set of families and the numbers of members of the sets of children they have. A more commercially relevant example is an indexical19 representation such as a security purchase and sale. Where, for example, an organisation’s trade is represented in its application as a security sale. But the same trade is represented in the counterparty’s application as a security purchase. It is only a sale or purchase relative to a party to the trade (and their application). The underlying trade whose existence these representations commit to (is made true by) is neither a sale or purchase in itself. Technology is also a common source of semantic divergence. The technology in which an application is implemented has a strong influence on how it is represented in the implementation. A database or programming language comes with its particular forms, and the implemented representation of the business objects must fit into these. The focus on business objects independent of the application and how it is implemented removes this influence – in other words, the model is technology independent. Experience with REV-ENG amply confirms the ubiquity of semantic divergence. Working applications are rarely straightforward ontological models – that is they have semantic divergences. Often it is the exigencies of constructing an application that 18 For example, the various chapters of Papazoglou, et al. (2000) Advances in object-oriented data modeling show markedly different designs for a standard car example. As its Chapter 10 Parent and Spaccapietra (2000) Database Integration notes there are a surprisingly wide variety of designs. 19 Indexicality is a common source of semantic divergence. It is where the truth of an expression (representation) depends the conditions of its utterance. A classical example is the expression “I am here” – which is usually true, but will refer to different people and places on different occasions. This is clearly a way in which we use language (representation) and not a way in which the world is. -- 12 of 43 -- The CEO Project: An Introduction Page 13 © Chris Partridge 2002. All rights reserved meets the enterprise’s requirements – and then maintaining it within a budget – that give rise to them. It is clear that the notions of semantic divergence and semantic heterogeneity overlap. What differentiates them is that semantic divergence assumes that there is a yardstick against which divergence can be measured – the underlying ontology – and so can measure this for a single application. Semantic heterogeneity merely notes differences in representation between applications. Hence, by itself, semantic divergence does not necessarily lead to semantic heterogeneity. If two applications are semantically divergent but have identical divergences, then they are not semantically heterogeneous. However, a close examination of the literature shows that it is recognised that dealing with semantic heterogeneity (typically semantically matching heterogeneous applications) requires some knowledge of the ontology (sometimes called ‘real world semantics’) and so, of necessity, semantic divergence. For example, (Vermeer and Apers 1996) notes “…schema integration techniques require either explicitly or implicitly that (the relationship) between the real-world semantics of the classes to be integrated is known.” 20 The REV-ENG experience is that much of the semantic heterogeneity in applications has its sources in differing semantic divergences. As the number of applications under analysis increases, the likelihood of this kind of semantic heterogeneity also increases. So, in practice, most ontological analysis projects have to deal with significant semantic divergence. A canonical scheme An ontological model can be seen as a canonical representation scheme. The notion of a canonical form comes from mathematics, where it is defined in terms of the general notion of a normalisation procedure, which consistently transforms objects (for example, matrices) to a canonical form. This enables one to determine whether different forms are equal relative to the normalisation procedure and its canonical form. In relational database modelling there is a well-known normalisation procedure for data that leads to a canonical form called the normal form. For computer applications, the ontological model can be seen as a semantic counterpart. Ontological analysis as normalisation One can see that ontological analysis is a kind of normalisation process for representations that leads to a canonical form in the shape of an ontological model. The normalisation can help to identify when the representations in different applications are of the same objects21 – and the ontological model is a direct representation of these objects. 20 This also notes how difficult this can be: “One of the central problems … is that the definition of relationships between local and imported data is far from trivial in a situation where information on the meaning of a remote schema is limited. … [I]n a federation of databases from multiple modelling contexts this may be surprisingly difficult.” 21 Or partially identical taking Armstrong’s approach (described in, for example, Armstrong (1997b) A world of states of affairs) to mereology. -- 13 of 43 -- The CEO Project: An Introduction Page 14 © Chris Partridge 2002. All rights reserved Canonicity and independence Many approaches to business modelling do not attempt to deliver either application or technology independence. This restricts the scope of the canonicity that they offer to the application and/or technology – a kind of local canonicity. By aiming for independence, the CEO will provide a more global canonicity (global relative to its domain – the enterprise). Benefits of a canonical modelling scheme The two main benefits of having such a scheme are firstly, that it provides a framework for re-use and generalisation and secondly, that it helps to facilitate inter- operability. Re-use is dependant upon recognising where opportunities for re-use exist. A canonical scheme allows one to recognise when the same business objects are involved and so that there is a possibility for re-use. Generalisation involves recognising when two or more types of business object share common general characteristics. In a canonical scheme similar types are represented in similar ways, making similarities easier to identify – and so facilitating generalisation. For interoperability, one needs to know when the representations in different applications refer to the same business object. A canonical scheme provides a basis for doing this by providing a framework within which the same business objects are represented in the same way in different applications’ business models. Canonical extendibility The CEO will provide a core framework around which applications can be built, often independently, extending the CEO to meet their needs. Many of these applications will have underlying business objects in common. To support general inter- operability, the ontological analysis (normalisation) process needs to work in a way that helps to ensure that the extensions are done in a consistent way: that the different extensions to the CEO independently represent the objects in the same way. Categorical ontology There is tradition that starts with Aristotle22 of not only ordering the types into a taxonomy but also explicitly including, at the top level, the major formal categories of entities (what can be called, more pompously, the types of existence). As a matter of principle, all the various lower level types fall under one or other of these top level headings. Following (Thomasson 1999), let’s call this a categorical approach. A number of philosophers have distinguished this categorical approach that attempts to provide an overarching structure from a more piecemeal approach that considers things on a case by case23 . They point out its advantages. For example, (Thomasson 1999) (on pp.115-6) notes a purely piecemeal ontology “can only provide a patchy view of what there is and a view that always risks arbitrariness and inconsistency.” 22 See Aristotle The categories. 23 As already noted this follows Thomasson (1999) Fiction and metaphysics (pp.115-6). Similar distinctions are made in: Williams (1966) Principles of empirical realism (see p.74) – see the distinction between analytic ontology and speculative cosmology is made and Ingarden (1964) Der Streit um die Existenz der Welt. 1. Existentialontologie (pp.21-53) – see the distinction between ontology and metaphysics. Similar points are made in the Introduction to Hoffman and Rosenkrantz (1994) Substance among other categories. -- 14 of 43 -- The CEO Project: An Introduction Page 15 © Chris Partridge 2002. All rights reserved and (on p.117) “Approaching ontological decisions globally avoids the dangers of inconsistency and false parsimony that may result from piecemeal ontology.”24 Computer science has picked up on the value of a categorical ontology. For example, John Sowa, in his latest book ((Sowa 2000) on p.51), states that “A choice of ontological categories is the first step in designing a database, a knowledge base or an object oriented approach.” Core enterprise ontology The ontology produced for the new accounting schema can be divided into a number of layers. At the top are the formal categories25 . Underneath this is the core enterprise ontology. A core ontology – as (Breuker, Valente et al. 1997) note – “contains the categories that define what a field is about.” Where a “field is a discipline, industry or area of practice that unifies many application domains …”. Determining the scope of core ontology, and, in particular, the boundary between the top and core ontology, is a practical matter – and is guided, as Breuker et al suggest, by how much a candidate category helps to provide a unifying structure. The key point is that given a ‘field’ such as accounting there are core categories that help to “define what [it] is about.” Epistemology There are two reasons why it is useful to introduce the notion of epistemology here. Firstly to clarify by contrast the notion of ontology and secondly because any new applications built using the CEO will need to have an epistemology built on top of their ontology. In philosophy, ontology and epistemology deal with two different questions, which result in two different ways of looking at and analysing the world. Ontology is concerned about what exists – whereas epistemology is concerned about what is (or can be) known by someone. For example, epistemology would attempt to explain how we can know about a particular type of thing, such as colours. Whereas ontology would be interested in what ontological type colours are. These two different approaches are both useful when specifying a system, particularly a computer application. A system (application) will make some ontological commitment – it will assume that certain things exist. These things are its ontology, which answers the question – what exists according to the system. The ontological model will represent this. A system will also have constraints on what it actually does (and can) know26 . These are described in its epistemology, which answers the question of what the system can (and must) know. In this context, an epistemology is always indexed to a knowing system. Of particular importance for operational applications is describing what it needs to know before it can do something. An epistemological model will represent this. Philosophical epistemology includes consideration of questions of belief, particularly the problems of false belief. Specifications of computer systems seem less concerned about these. 24 Similar points are made in, for example, Collingwood (1940) An essay on metaphysics and Körner (1970) Categorial frameworks. 25 In Partridge (1996) Business Objects: Re - Engineering for re - use this is called the framework level and an example of this for IT ontological analysis is given on pp. 276-8. 26 In the case of a computer application this system may be a network of applications, each with its own constraints upon what it can know. -- 15 of 43 -- The CEO Project: An Introduction Page 16 © Chris Partridge 2002. All rights reserved One can regard the epistemology as looking at the world from the perspective of the system and what it knows and the ontology as standing back and describing the world that the system commits to from a perspective outside it. These two are interdependent. They deal with the same world, and mostly with the same things in that world. However, their different goals mean that they paint different perspectives of these – as the following examples show. Examples Let us assume, simplistically, that all humans are either male or female and that we are looking at a system that records humans’ details including their gender. Then this system is ontologically committed to the existence of male and female types, which are sub-types of human and completely partition it. This is its ontology. However, we cannot guarantee that the system will always know a person’s gender – so it has to deal with cases where it does not know the gender. So within the system’s epistemology not all humans will be partitioned into male or female sub-types – in other words, within the epistemology the partitioning is incomplete. This gives us different, but equally valid, ways of categorising the world, illustrating how the approaches’ different purposes can lead to different results. Epistemology’s purpose lines up quite neatly with one of the key requirements in specifying a computer application, clarifying what it must know and what it does not need to know. This makes documenting the epistemology an essential element of the specification of a system – though it is not usually called given such a grand name. To see this, consider an insurance company that sells various types of policies. For its actuarial calculations, the company needs to know and so asks all its policyholders whether they are married and records the results. For its joint policyholders, it also needs to know, and so asks, whether they are married to each other and if they are, this marriage relationship is recorded. For its sole policyholders it does not need to know this information – so it does not ask for or record it. The system is ontologically committed to the existence of persons and their married states. It is also committed to the fact that persons in married states have a marriage relationship with each other: this is what being in a married state means. In contrast, from the company’s epistemic perspective, knowing someone is married does not mean knowing their spouse and marriage relationship. This is because for sole policyholders, the company can know that they are married, but not know who their spouses are and so cannot know their marriage relationships. Note that it may ‘know’ their spouses – because they are also policyholders – but not know that they are the spouses. In current practice, the epistemic perspective plays a more prominent role in computer specifications because the current state of database technology means that the epistemology (and not the ontology) is reflected more directly in a company’s database. In this example, the insurance company’s database needs to be able to record persons that could be in a married state without having to record them having a marriage relationship. The fact that persons in a married state always have a marriage relationship cannot be recorded. This is why the use of the terms ‘mandatory’ and ‘optional’ for attribute and relations in database contexts are usually from an epistemic (not an ontic) perspective. -- 16 of 43 -- The CEO Project: An Introduction Page 17 © Chris Partridge 2002. All rights reserved Linking ontology to epistemology It is important to understand how the ontology and epistemology link. One way of analysing this is to widen the scope of the ontology to include the system that is the subject of the epistemology (though this can pose some delicate problems and needs to be done carefully). Consider the first example. It may be tempting to regard the epistemological model as representing epistemic types that deal with known instances – as perforce only these are instantiated in the model. This would introduce new epistemic sub-types of human: known-male and known-female – and possibly unknown-gender. However, it makes more sense to say that the system ‘knows’ the ontological types male and female, though it does not know all their instances – it may even know an instance of human, know it is male or female, but not know which. This captures our unreflective view of our own epistemology, which regards its male or female sub- types as ontological; in other words, as referring to all males and females not just the ones we know. Also, it avoids the possibility of an endless regress. Opting for epistemic ‘known’ sub-types would introduce the possibility of an endless regress As it is possible for a system to know whether it knows, one would need to also introduce ‘known-known’ and ‘known-known-known’ sub-types and so on. Under the ‘ontological types’ option, the ontology would capture the epistemology by explicitly recognising the system and its knowing relationship with the male and female sub-types and their instances. It would also recognise that only some of the gender types’ instantiation relations are known. So, for example, instances of human that are not epistemically classified by gender (in other words, whose gender is not known) would be marked in the ontology by not having a known relation between the system and the instantiation relation. This explains why the epistemic partition is incomplete – it is only representing the known instantiation relations. In this simple example, the epistemological perspective can be seen as a filtered view of the ontology – only showing what is known. This filtering led to the difference in structure. From this brief outline it should be clear that specifications for enterprise applications schemes need both an ontology and an epistemology. Applications sometimes need to be able to record that they know someone, who has a gender, but they do not know which one. Insurance companies may need to know that if their policyholders are married that they have a married relationship with someone else – even if they do not know who the person is. CEO sample ontologies and epistemology All of the best of breed ontologies selected are an amalgam of ontology and epistemology. There are reasons for this. The AI based ontologies take a conceptual view of ontology and within this perspective, no distinction is made between ontology and epistemology. The data modelling based ‘ontology’ is meant to represent data that will be stored in an operational database27 – which, of necessity involves epistemology. One element of the ontological analysis will be to filter out these epistemological aspects. 27 This involves an element of equivocation as the model is at one time a representation of the business and another a representation of the data that, in turn, represents the business. However, this kind of equivocation is endemic in data modelling – and elsewhere. In philosophy it is often called a use mention confusion – as it confuses the use of a representation with mentioning it. -- 17 of 43 -- The CEO Project: An Introduction Page 18 © Chris Partridge 2002. All rights reserved The nature for ontological analysis It may appear that having clarified what an ontology is, the next step is to work with the relevant experts to organise what they know about their domain: in the case of the CEO, with enterprise experts, to organise what they know about the enterprise. However, experience in building enterprise models, and more generally in specifying application requirements, shows that this is not a successful approach. Experts typically cannot articulate what it is they are working with28 . To understand this one must distinguish between two types of knowledge: know-how and what I shall call, know-what. Know-how is, traditionally, what experts have, otherwise they could not do their job. Know-involves the ability to articulate what the entities involved actually are. Initially, the fact that experts do not have know-what may seem strange. But there is an example that we are all familiar with. We are experts in the use of our mother tongue. Our ability to apply subtle grammatical and syntactical principles is astonishing. We have language know-how. However, our inability, unaided, to articulate the principles that we are using is clear. We do not have know-what. (Strawson 1992) 29 using a similar example, points out that this shows knowing-how in no way implies knowing-what Given this situation there is a clear need for a process to enable the articulation of the know-what. (Strawson 1992) points out that philosophical, metaphysical analysis is a traditional way to get a systematic representation of our know-how – our know-what. An important foundation for this analysis is an understanding of why there is this gap between know-how and know-what. We develop this by looking at why it exists30 for the institutional and social facts that are the subjects of most enterprise data. Socially constructed business objects Institutional and social facts are of a different kind than ordinary everyday physical objects, such as trees and stones. Physical objects have an existence independent of us. Whereas most institutional and social facts (which includes most business objects) are dependent upon us – and often seem to be constructed and maintained by us. This makes it even more difficult to understand why experts cannot articulate what these are. Money is a good example of an institutional and social fact. People have used many things as money, including cowrie shells. What makes these money is that those 28 There are examples of this inarticulacy in Partridge and Stefanova (2001) A Synthesis of State of the Art Enterprise Ontologies and Partridge (2002d) STPO - A Synthesis of a TOVE Persons Ontology (forthcoming) and Partridge (2002a) What is pump facility PF101?. 29 Strawson (on pp. 5-7 of [Strawson 1992]) makes a similar point using the example of the first Castilian grammar being presented to Queen Isabella of Castile (in the Ninth Century). She asked what use it was, because “in a sense [Castilians] knew it already. … though in a sense they knew the grammar …, there was another sense in which they did not know it.” He draws “ the general moral that being able to do something … is very different from being able to say how it is done; and that it by no means implies the latter.” Noting that “In contrast with the ease and accuracy of our use are the stuttering and blundering which characterise our first attempts to describe and explain our use.” Interestingly, he goes on to suggest that “the philosopher labours to produce a systematic account of the general conceptual structure of which our daily practice shows us to have a tacit and unconscious mastery.” 30 That it exists is acknowledged - see, for example, Searle (1995) The construction of social reality and Gilbert (1992) On social facts. -- 18 of 43 -- The CEO Project: An Introduction Page 19 © Chris Partridge 2002. All rights reserved people accept them as such. Cowrie shells are certainly not intrinsically money, independently of the humans that use them as such. Similar things are true of languages and social institutions, such as marriage. The same is not true (at least not in the same way) of trees and stones. (Searle 1995) analyses this difference and describes these people-dependent objects as socially constructed and calls them human institutions. Furthermore, the rules that characterise human institutions (to use Searle’s name) work in a different way from the rules that characterise physical objects. Physical objects have rules (laws) that govern their behaviour, but cannot be said to know the rules. Stones do not have to learn the rules of gravity before they fall – nor can they decide, once they know the rules, that they do not want to follow them. Whereas people have to learn the rules that govern their human institutions. This can be quite arduous, as, for example, when someone has to learn a new language. It is also possible (in many cases quite easy) to ‘disobey’ the rules. For example, fluent language speakers can and do choose to make deliberate grammatical mistakes. Most people are comfortable with notion that the conceptual structures of business and law that underlie the enterprise are socially constructed artefacts. However, with this notion often comes an assumption that is not so well warranted which is relevant to our topic: that these artefacts are solely the result of people following rules (even constituted by the rule following). This would imply that following the rules only involves consulting the rules in their heads. Ontological analysis would then involve examining these rules to develop a more precise picture – for example, by interviewing experts or training them in introspection. Though there are elements of truth in this assumption, I shall argue that it is mistaken in the case of analysing a core ontology and that a different method of analysis is more appropriate. Following rules The assumption seems, on the face of it, reasonable. As already noted, language is an archetypal example of a human institution. Consider someone who is learning a new language. When they try to speak, they have to laboriously consult the rules that they have been taught and are conscious of trying to follow them. Things are less clear for children learning their mother tongue. However, this can be explained, as Chomsky does in his account of Universal Grammar (Chomsky 1975). He reckons a child is able to learn grammar because he or she is already innately in possession of the rules of a universal grammar, though these are unconscious. Closer examination of specific cases shows that most rule following is unconscious. When someone has learnt a language properly, they are no longer conscious of consulting and following its rules. Similarly, practicing business people and lawyers are typically not conscious of the rules they are following. The examination also reveals deeper problems with the rule-following account – it does not seem to fit the obvious facts. As Searle points out (on p. 127): “the structure of human institutions is a structure of constitutive rules ... the people who are participating in the institutions are typically not conscious of these rules; often they have false beliefs about the nature of the institution, and even the very people who created the institution may be unaware of its structure.”31 31 Searle (1995) The construction of social reality, p. 127. Gilbert (1992) On social facts makes a similar point. -- 19 of 43 -- The CEO Project: An Introduction Page 20 © Chris Partridge 2002. All rights reserved Even when the beliefs are true, they are often inadequate by themselves for their purpose. As every system designer knows, experts typically cannot articulate rules to a sufficient level of formality and precision. Even though they have no problem in actually undertaking the tasks precisely enough. If we know these rules and are following them, it seems strange that we can so regularly have false beliefs about them. Particularly when this seems to have no correlation with our ability to follow them correctly. It also seems strange that we cannot articulate the rules to a level of accuracy that we must know to be able to follow them properly. The problem is in the assumption that we are always following rules. Searle articulates32 the issue as a question about the causal role of the rules, which neatly distinguishes between the two extremes in the ways in which rules operate. Are there rules in our heads that are the cause of us following the rules – in other words, are the rules representations which we consult and follow? Or, at the other extreme, do these rule representations have no direct causal role – merely providing a description of the actions we take? In other words, we do not ‘follow’ the rule representations. Neither extreme seems to fit all the evidence. As noted earlier, human institutions are clearly not completely governed by rules in the way physical objects are. But, on the other hand, neither are they completely subject to forms of rule following. A number of philosophers33 (including Searle) have suggested that our more conscious rule following is grounded in natural propensities that operate at the level of neurophysiological (non-intentional, non-representational) processes. This implies that the closer the rules are to the foundations, the less rule following is involved. Nature of ontological analysis Irrespective of the chosen explanation, these facts have a clear implication for the process of analysis involved in building a core ontological model. Given the level of false and inaccurate knowledge there is of the rules that experts are capable of articulating, and that, in many cases, they are incapable of articulating the rules, it does not make sense to based an analysis methodology on the presumption that they have a sufficiently complete and accurate knowledge of the rules. While interviewing experts may play a part in the analysis, it will never give a complete picture. And experts’ claims will need to be examined in the light of what actually happens, what people and organisations actually do. CEO’s requirements for a core ontology The CEO’s requirements will guide its choice of ontology and its use. Because it makes the relevancy of the requirements clearer, we consider the specific CEO requirements first, and then set them in context, by looking at the general requirements for a core ontology. 32 Gilbert (1992) On social facts p.127-8. 33 Wittgenstein (1953) Philosophical investigations introduced the question of how we ‘follow’ rules and discussed natural dispositions. Kripke (1982) Wittgenstein on rules and private language revived the discussion more recently and it is now a lively topic. See also, for example, this point in Wright (1987) Realism, meaning, and truth, p.28 “… the path to understanding exploits certain natural propensities which we have, propensities to react and judge in particular ways. The concepts which we ‘exhibit’ by what we count as correct, or incorrect, use of a term need not be salient to a witness who is, if I m
BORO Research
The CEO Project:
An Introduction
30 November 2002Published in LOA (LADSEB-CNR), Technical Report 07/02, December 2002, Padova, ItalyLOA (LADSEB-CNR), Technical Report 07/02, December 2002, Padova, Italy
Overview
This is, in essence, the project initiation paper for the CEO Project. Its main concern is explaining the project’s aims, how it intends to achieve them and the methodological framework within which the project will work. It explains the origins, conception and motivation for the project and gives an outline of the management framework for the project, in particular the first synthesis stage. It clarifies the terms of the art and describes the nature of ontological analysis. It also characterises the requirements that shape it and the meta-ontological choices and analytic styles that underlie it. Finally, it describes the potential applications and the next steps.
