Showing posts with label research activity data. Show all posts
Showing posts with label research activity data. Show all posts

Monday, 9 August 2010

Uncovering hidden connections in Research Activity Data

I recently submited an article to Sconul Focus which, I hope, will be published in a few months time. The topic of the article was data harvesting and aggregation. A simple and (hopefuly) easy to read explanation of a rather complex topic. I used our entity registry as an example and described how it harvests data from University and external sources, how it converts sources into RDF format so data can be aggregated and finally how it uncovers hidden connections. Regarding this last step, uncovering hidden connections, I thought this was a very interesting and value adding process, so I will blog about it here.

When we collect data from different sources, there is a chance that some of these data are interconnected. Different sources may contain replicated data, e.g. a researcher's profile can be found on his college and department's websites. Different sources may contain data that complement each other. For example a researcher's profile on his college website and a list of his projects and publications on his departmental website. Originally, these webpages do not include links to each other therefore unless we know about the other source we will not get a complete profile for that researcher. Add more sources, external ones too (e.g., funders' websites containing information about grants), and we will get information about researchers which is spread all over the place but disconnected.

Uncovering hidden connections means making connections between data in different sources more evident: as in adding links between sources which point to every each of them. But how do we create those links? Here a non-technical introduction:

For example if we get Prof Francis Matthew Kellner's profile in source 1 and Prof Francis Matthew Kellner's research interests and publications in source 2 we can establish with some level of certainty that these two sources refer to the same person. Therefore we can connect the data in these two sources and build a completer profile with Prof Kellner's profile, research interests and publications.

There are, however, other cases which are not so straightforward, where names are similar but we cannot be sure they belong to the same person. For example if we get Prof Francis Kellner’s biography in source 1, Prof F. M. Kellner listed as Principal Investigator on a project in source 2 and Francis M. Kellner as author in publications in source 3. How do we know if these data belong to the same person?

For cases like these we have developed a ‘same-as’ process.

‘Same-as’ has a set of rules which use information such as people’s first name and surname, researchers’ affiliation and email. Depending on the availability of information and whether the sets of data match, ‘same-as’ determines if two or more records belong to the same person or not. If the records belong to the same person ‘same-as’ merges the records. If the information available is not enough to do the matching, or if the data do not match, ‘same-as’ will keep the records separately.

The following is the logic used. This is a technical-ishh explanation written by Anusha.

Search for people with the same last name, who are part of a group of sources, e.g. sources belonging to the social sciences (we group sources together, based on likelihood of information overlapping)

Case 1
Match people only if each person has at least the fields below and they all match.
  • first name (not just initials)
  • last name
  • affiliation
  • source(s)
If the source is a trusted source:
  • staff_id
Note: Subset of the firstname will be matched : Example John P M, John P, John
Create a person superset and add all of their info

Case 2
For people with the information in atleast these follwing fields, with all of them matching:
  • initials
  • last name
  • affiliation
  • source
Create a new person superset and add them to that. Treat them as a separate person and do not add them to the person above
If the person matches the information above, suggest a connection.

Case 3
For people with the information in atleast these follwing fields, with all of them matching:
  • initials
  • last name
  • source
Create a new person superset and add them to that. Treat them as a separate person and do not add them to the person above.
If the person matches the information above, suggest a connection.
If the source is a trusted source (data harvested from within oxford), display them in the browse / search results pages, else do not display them.


The ‘same-as’ process focuses on ‘people’ entities. Projects, publications, funders and academic units, usually have fixed (or standard) names which are used consistently across sources. However, names of people are frequently written in different ways, depending on the contexts.

So now, you can imagine, everytime we add more data to the registry, we pass these data through the same-as process to see if there are any hidden connections to the data we already have. Or put it in a different way, everytime we add a new source, we do not only add their data but the connections we find with same-as, which were probably unknown or at least not evident in the original sources.

If you want to know more you can wait for the Sconul Focus paper or if you are impatient e-mail us.

Friday, 2 October 2009

Institutional Repositories: acceptance and adoption

I have been reading some literature about institutional (digital) repositories. I am interested in the human and social issues surrounding their development and implementation as well as their embedding. I think this literature is very relevant to BRII as it reports experiences in similar implementations (not necessarily technically but conceptually) in similar kinds of institutions.

Technically speaking BRII is not building a digital repository but a Research Information Infrastructure (so a bit much broader in scope.) However the issues, technical and non-technical, surrounding implementation of institutional (digital) repositories apply. The reasons are:
  • BRII has an institutional scope and it is situated within an academic context. It aims to collect information about research in all academic areas and its target audience is everyone within the University who carries out research-related activities.
  • BRII deals with Research Activity Data: information about research. So, not information used and generated by research (as in datasets and publications) but descriptions of research (people and activities.) This kind of information is essential to facilitate scholarly communication.
  • BRII core users are Researchers and Faculty: as content contributors of research descriptions and as users of the information deposited in the infrastructure.

Being now in the "development stage" of the project we are starting to think on how to make our products more relevant to our core users, so they understand them and use them. Of course to make users understand the Research Information Infrastructure we first need to understand our users. Our work in BRII should not be techocentric only but should expand to reaching out to our users, speaking their own language and observing them in their own working spaces. This can be a very difficult task as we are dealing with heterogeneous groups of academics and administrators, each one with different needs and perspectives (as per their research culture.)

Understanding our users is a crucial first step to achieving acceptance and adoption of the RII.

Research in the area of Institutional Repositories field has reported some issues concerning this.

In a study of implementation of an Institutional Repository (IR) at the University of Rochester, Foster and Gibbons (2005) report a "misalignment between the benefits and services of an IR with the actual needs and desires of faculty." They state that the benefits of Institutional Repositories are attractive only to the institutions which host the repositories. This is because Institutions (including Librarians and IT developers) see IRs as a way to facilitate access, reduce costs and improve efficiencies at collecting and managing information that is produced in the Institution. On the other hand academics see themselves as belonging to research communities (not institutions) and are not concerned about the proccesses of managing all that data they produce. This is, perhaps, because they see the problem as affecting others and not themselves. (Colbertt, 2009)

At the centre of this all is the Scholarly Communication crisis. In 2002 the Scholarly Publishing and Academic Resources Coalition (SPARC) (part of the association of Research Libraries (ARL)) stated that Institutional Repositories could help combat dissatisfactoin with the "monopolistic effects of the traditional and still pervasive journal publishing system" (Crow (2002) quoted in Maness et al (2008)). Rieger (2008) states that to combat this crisis, academic institutions aim at "reducing costs of producing and acquiring publications and gaining control of processes from commercial publishers." Seen from this angle, institutional repositories are good and sensible solutions. Attention therefore is devoted to technical efficiencies and control of the information produced within the institution. A consequence of this is a bias towards computer and librarian approaches to developing IRs, which neglect perspectives from relevant groups’ interests.

From BRIIs Stakeholder analysis we have learned that technical efficiencies, metadata, repositories are terms which are meaningless to academics. This concurs with Foster and Gibbons (2005) view. Seen from this technocentric perspective IRs do not provide benefits to academics. In addition having institutional labels suggests researchers that the IRs will support the needs of the institution and not their own individual needs. Academics are indeed a different crowd:
  • Academics think in terms of reading, researching, writing and disseminating.
  • Academics have strong ties with the people interested in their own field of research, or with whom they are interested to collaborate. Their geographical location is not important to them.
  • Academics are interested in a relatively small subset of research information, that one of their own research field. They have acquired skills to search for and use that information for their own work.
BRII Stakeholder Analysis (Loureiro-Koechlin, 2009)

On the other hand implementers of repositories (IT and library practitioners) possess an institutional perspective, they are interested in developing new forms of scholarly communication, control costs and improve data manipulation.

An example of institutional perspective is the mapping of people and information under the umbrella of their institutions and departments. This is an obvious mismatch as researchers see themselves as belonging to their research communities. Other issues connected to the particular ways in which research is carried out (technical and ethical) also accentuate these differences. One example of this is the differences of methodologies and nature of data collected by natural scientists and social scientists. Carusi and Jirotka (2009) report ethical issues of archiving qualitative data in a digital format, such us consent, anonymity and privacy…They state that reusing research data for other purposes will go against research participants wishes and therefore researchers ideals.

Misalignments like the above are reasons why academics do not feel institutional repositories are relevant to their work. Therefore as Foster and Gibbons (2005) report there is a need to get to know researchers in their own space so we can develop tools which are useful to them as well as to the institution. This improves the chances of IRs to being accepted and adopted by their users.

One example is Maness et al's (2008) study on needs and goals of institutional repository. It reports results which contradicted initial assumptions by IR designers and decision makers. While designers and decision makers assumed users wanted an open-access archive for research outputs, the study found out that in reality users wanted a network to share learning and teaching materials, where collaborators could be identified and where their research could be promoted to institutional colleagues.

Carusi and Jirotka (2009) From data archive to ethical labyrinth. Qualitative Research.
Corbett (2009) The Crisis in Scholarly Communication, Part I: Understanding the Issues and Engaging Your Faculty. Technical Services Quarterly, vol 26 p.125-134.
Foster and Gibbons (2005) Understanding faculty to improve content recruitment for institutional repositories. D-Lib Magazine, Jan 2005.
Maness, Miaskiewickz and Sumner (2008) Using Personas to Understand the Needs and Goals of Institutional Repository Users. D-Lib Magazine, Sept-Oct 2008.
Rieger (2008) Opening up institutional repositories: Social Constitution of Innovation in Scholarly communication. Journal of Electronic Publishing.

Thursday, 17 September 2009

Data, ideas and more

Yesterday, I met with Luis Martinez Uribe, Digital Repositories Research Coordinator, to talk about research data. Luis is working in the The Embedding Institutional Data Curation Services in Research (EIDCSR) project which looks at research data management and curation challenges. Previously he worked in the Scoping digital repository services for research data management project where he collected very interesting data which forms the foundation for the EIDCSR and hopefully many more similar projects.

Luis and I talked about overlaps in both projects and on how some common themes arised in both our data analyses. Luis has collected information about the processes and needs of researchers regarding the creation and use of research data. These are data that researchers create and use as part of their research activities (e.g., images of the heart. ) He is looking at developing processes for creating metadata and preserving both metadata and research data across time. In this form data can be reused by multiple researchers in multiple ways.

In BRII I have collected information about the nature of research activities across the Scientific - Human spectrum and from the academic, administrative and strategist perspectives and how these activities are reflected in public data such as websites. Part of the information collected reveals the reasons for sharing Rearch Activity Data (e.g., improve visibility) which are the reasons BRII wants to support and enhance across the University and beyond. Of course data have also revealed that there are reasons against making Rearch Activity Data publicly available (e.g. confidentiality issues) as well as technical difficulties which make sharing a complex task.

Anyhow, we found that there were common issues arising from the creation, use and sharing of research related data: Research Data (EIDCSR) and Research Activity Data (BRII) particularly within the context of institutional initiatives (as opposed to individual efforts). We think that these issues deserve some attention as they are influential in the implementation and acceptance of efforts like EIDCSR and BRII by their core users: Researchers.

...So we are outlining a paper which will explore those issues and let's see what happens.

Friday, 27 March 2009

Some thoughts about Data

One of the interviews I conducted this week left me thinking on the issues of transformation of data and its different meanings. I was talking to this lady about her work with research activity data. She explained me how she collects information from various sources, enters that information in a spreadsheet and accommodates it according to her needs. This accommodation involves the organization of data in columns, the correction of errors, filtering records and adding new ones from other sources. All this transformed data will be entered in a research portal.

I asked my interviewee whether she could contribute some of the data she has so far. She said I could get all that information from the same sources she used, that all of those sources were public. And that made me think…. On the one hand, I thought she was right. All the information she had came from other sources which we can all access. She had not created new data but just worked on existing data. However, on the other hand the sets of data she had been working on represented new pieces of information. The work she carried out on that data transformed it in new data. Data+Work=NewData. So I guess new arrangements of data provide new meanings to that information.

A simple example: You can access an international online database and download a list of publications on economy. You enter that list in a spreadsheet and filter the publications belonging to a particular Oxford University author. Then you attach that list to the author’s profile and you get his bibliography, all his publications since before he joined Oxford. You can do that with all economists in Oxford and you will get a number of bibliographies from Oxford economists. This list will correspond to the produce of Oxford economists across their careers.

You can also use that list to filter all the publications from authors whose affiliations are Oxford University, current staff, or staff who has already left, but who produced the publications while they were working in Oxford. This second list is the produce of Oxford economists in Oxford.

Both lists come from the same source, and possibly they contain the same set of fields, however they represent different things.

When I did my PhD I came across a book called “Information, systems and information systems: making sense of the field” by Peter Checkland and Sue Holwell (1998) a must read if you are in the Information Systems field! This book is about information systems, their creation and relation to IT. In chapter four they discuss the concepts of Data, Information and Knowledge, and they introduce the concept of Capta. These concepts may help you to understand all these processes of transformation of data and how they can acquire different meanings. For Checkland and Holwell (1998) data represents all these masses of facts, observations and concepts that exist in the Universe. Once data are captured as part of an information system, a conversation or any kind of interaction they become Capta. Capta therefore are a subset of data which have been selected through a purposeful process, i.e., according to a criterion which fits a particular purpose. Capta are transformed into Information when they are given meaning and context by their interpreters. Because they depend on interpretation, a subjective process, information can have different meanings to different people. Finally, large structures of information form Knowledge.

Now, how can I explain this in the context of the BRII project?
Well, all these processes of transformation of research activity data into capta and into information happen all over the University. People acquire research activity data from different internal or external sources and transform them according to their needs. New people may use these transformed data and give them new meanings, again according to their new contexts. This seems like a mess, but
  • BRII will sit in the middle, facilitating these processes.
  • BRII will extract capta from data and store it in the Research Information Infrastructure (RII).
Using Checkland and Holwells’s concepts, we can define Research Activity data as capta. Data selected from vast sources which represents and describes only research activities. BRII will ignore data which does not fit this criterion. So the RII will be a container of capta in the sense that it will only host research activity data.
  • If we see this from a different angle, within the universe of the RII and call its content data again, I can say that BRII will provide the means to reuse that data and transform it into capta, capta for every system or individual who access the RII looking for information. (Different purposes and different contexts.)
As explained in the example above, most sets of existing data are data which have been filtered and worked on according to criteria which depend on the contexts and purposes of their owners. Another example, the list of researchers in a departmental website is a subset of researchers of the University. This subset was selected by checking on the affiliation of each researcher to the department or his/her work within a research group or project within the department. The same researcher may appear in another website as he/she is involved in other research activities. However, this researcher does not appear in a thematic website as his/her interests are different. Ideally the RII will hold a list of all researchers in Oxford and by accessing the RII people will be able to extract these subsets of data. This data will acquire different meanings depending on the criteria used for its extraction and on the place and way it is shown.
  • The RII will be a tool to give meaning to huge, disparate, disconnected, complex set of Research Activity data.
The RII will be a big Information System supporting and allowing other systems to exist on top of it. The RII itself and the web services built on top of it will allow the transformation of data into capta and information by its users.

What am I in this context?
I am the person who is looking for these sources of data and capta and tries to understand what they mean to their users and what new meanings can be obtained from them in the future.

Some issues
How do we know the purpose of capta, capta which has been adapted by some people from other sources of capta also within the University. Should the RII store only sets of raw data and allow its users to transform them? Should we use all these sources, data and capta? How?
Again, the answer lies on the semantic web. Semantic web technologies allow the labelling of data with labels which are meaningful to peolple and to computers, that is, tagging data with meaning. From the context of the RII these labels convert data into capta. From the context of the services accessing the RII tags are the means through which users can convert data into capta for their own needs.

Friday, 27 February 2009

Connecting stakeholders and data

The first stage in the BRII project is to carry out a stakeholder analysis. A stakeholder analysis is a process that involves the identification of the project's stakeholders and an assessment of their interests and the ways those interests could affect or be affected by the project. Stakeholders could be people or organisations and their interests could be in favour or against the project.

As the BRII project is about gathering and using Research Activity data, our stakeholders will be people or entities who produce or use that data. For example, researchers, project managers, research facilitators and administrative staff. Also units, departments, institutes, colleges, where research is carried out.

The assessment of the stakeholders' interests will give us some light on the kinds of research activity data they produce and use for work. It should also be useful to investigate other related needs that are not fulfilled at the moment, but which could be covered by the project's outcomes. So, it is not only knowing who, but also what they do and what they would like to do. But more importantly it involves getting that precious data from them which we can use to build the Research Information Infrastructure.

In a previous post I wrote about finding potential stakeholders within Oxford complex organisation. I have also written a bit about the research management and research activity data we need from them.
In this post I would like to put these two concepts together.

I would like first to make emphasis on the importance of being able to connect stakeholders and their data between them. Research is an activity which is carried out within a research community. People from that community need to communicate between them, exchange ideas, collaborate, compare their research, assess the uniqueness of their contributions, etc, etc. One way of doing this is by accessing information about research that each other have, information which could reveal interesting opportunities, connections and collaborations that are happening around them.

I have an example that could illustrate this. During this "making sense of chaos" stage (which is still going on) I have been able to identify some links between units in MedSci and other departments in other divisions as well as with external bodies. I call them paths of cross disciplinary research collaboration. In the figure on the right two paths are shown:

1) Department of Cardiovascular Medicine - Centre for Research Excellence CRE (consists of research groups from across the University) - British Heart Foundation BHF (external funding body)

CRE is funded by BHF and promote cardiovascular research from all areas within Oxford.

2) ComLab (in MPLS) - Computational Linguistics Group (interdisciplinary group) - Phonetics Laboratory (Humanities division) - Babylab (MedSci).

I have come to know about some projects and collaborations happening across this path.

Like these two examples there could be many more paths out there which could be made visible through their data and BRII. Samples of paths like the above and other areas in MedSci will be good examples as they represent the heterogeneity of research in Oxford. Heterogeneous data will help BRII create richer ontologies which represent a variety of stakeholders and research their activity data in Oxford. Similarly, getting documented needs from a variety of stakeholders will help us to design web services which can demonstrate different uses.

BRII's expected output is a pilot of the research information infrastructure. Creating a pilot means creating a piece of software and data that we could use to demonstrate that the BRII strategy is feasible and useful. Having stakeholders representing different areas within the University (but mainly from Medical Sciences*) will help us to build a pilot which accounts for different kinds of data and needs (not all of them of course). That is, a pilot which can demonstrate a small set of services which account for diverse needs to serve an heterogeneous organisation.

* The Medical Sciences division is BRII's main stakeholder.

Monday, 16 February 2009

What is Research Management Data?

BRII's aim is to develop an infrastructure that captures research management information from different sources, which are published or which are classified as appropriate for open publication, and allow their sharing or dissemination in different contexts. As I mentioned in my previous post, I am planning a stakeholder analysis (which I hope to start soon!) to identify stakeholders and their needs in relation to research management data. Some stakeholders will also be sources of information. For example, BRII could use information in departmental websites, databases, or other kinds of documentation.

How do I know what information is relevant? BRII needs every sort of information that comes under the umbrella title of research management data.
So, then, what is research management data? In general terms, Research Management data could be defined as the metadata about research. This is not the data collected, processed or generated by the research process itself. It is the data that identifies each research activity, its components, elements, members and relationships. It tells you the who, what, when, where (and why?) of research in an institution.
Developing this definition more gets difficult in Oxford due to its organisational complexity. Different divisions, departments, units would have different perspectives which could depend on the aspects of research and/or the research fields they are involved with. These differences are also intensified by their autonomy, having independent business processes each. Also each unit will own different kinds of information, which are hosted in different sources and in different formats. Some could be written on yellow post-its stuck on somebody's whiteboard!

Anyway, the uses of research management data could be enormous. Two examples: external funding entities could use these data to assess research proposals and funded research. The data could also be used to produce reports for the Research Assessment Exercise (R.A.E.) or the future Research Excellence Framework (R.E.F.).

Besides my work at identifying potential stakeholder names, I also need to identify the kinds of information that I will need from them which equals to: the kinds of information that are relevant and available and that would be useful for our potential users. I am sketching a sort of map of research management data, from what I know about research and from what I have read so far since I took this post. What I have so far is the following list:

Research activity data: purpose or research, research field, research outcomes, website, institutional links, connections/collaborations with other research activities, Bids/Proposals, plans, research progress reports, project evaluation reports, other research proposals generated, people&roles, facilities, resources, technology, etc.

Social Data: people based (researchers, postsdocs, admin, etc) CVs, Bios, Publications, Interests, websites, etc.

Financial data: information about funding bodies, grants, budgets, time scales, extensions, overheads, etc.

Spatial/Geographical data*: where are people, projects located, laboratories, facilities, etc who works near who, who has access to facilities, etc.

Research Logs - things that happen in a project, issues, rules, changes in strategy --> don't think this would be classified as appropriate for open publication though

So one of my objectives will be to ask for these kinds of content in whatever format they are, and also, of course, I will need to ask stakeholders if they think there is something I should add to this list.

* Spatial/geographical data is a potential area for collaboration with the EREWHON project

-------------------
Note: since I wrote this post the BRII team decided to use the term Research Activity data to describe data that will be harvested. Data from interviews from the stakeholder analysis and feedback from other sources shows that Research Activity data represents the kind of information which is usually made available online and which is of interest for most of our stakeholders.