Usage Data – Self Plagiarism is Style

Dumping the OPAC #2 – Usage Data

File this under “things I wished I’d started doing a long time ago”, but we began collecting usage data from Summon a few months back, using a modified version of Matthew Reidsma’s Summon-Stats script. In that time, we’ve built up a log of just over 310,000 user transactions from Summon searches for around 160,000 distinct items[1].
It’s early days yet and we’d need to collate more data before we could launch this as a truly viable service, but we’ve got enough usage data to start driving “people who clicked on this item, also clicked on these…” style suggestions (based on the same methodology we’ve been using since late 2005 on our OPAC):

Although the suggestions seem to mostly be relevant and on-topic even at this this early stage, given more usage data, they’d become increasingly refined and improved.
I guess conceptually this is similar to the Ex Libris bX service, but it also exposes items that wouldn’t necessarily be tracked via link resolver logs (e.g. books, web sites, off-air TV recordings, items in the repository, etc).

[1] – that represents just 0.12% of the items in our Summon instance

The Magic Microphone and the Horrible Headphones

A while back, I had a daft idea to try and use all of the library circulation and graduate attainment data to pinpoint which books were the ones most likely to be borrowed by high achieving students, as I thought it might be fun to highlight those on the OPAC — after all, if you had the choice of two books on a topic, why not go for the one that the brightest students had borrowed the most?
Anyway, I came up with a formula to calculate a percentage score for each item, based on the attainment of the students who borrowed it and the overall distribution of all graduates, where 100% meant that only students who achieve a first class honours degree borrowed the item. I then discarded items borrowed by less than 100 students and held my breath to see which scholarly tome achieve the highest score. What I hadn’t expected to discover was that we appear to have some magic microphones, capable of bestowing high marks to all who borrow them…

Quite why these Shure SM58 microphones seem to mostly be borrowed by students who gain the best honours degrees, I can’t say!
On the other hand, the item with the lowest score turned out to be a set of headphones, borrowed mostly by students who achieved a 2:2 or a third…

The more superstitious amongst you will be glad to hear that the headphones have since been weeded … and, for all I know, buried in an unmarked grave at midnight 😀
For those of you who are interested in such things, here are the books that scored highest:

Neuroanatomy : an illustrated colour text (80.08%)
Audio-vision : sound on screen (78.73%)
The foundations of social research (78.40%)
The computer music tutorial (78.38%)
Critical discourse analysis : the critical study of language (78.24%)
Understanding the self (78.09%)
Doing qualitative research : a practical handbook (78.01%)
Personality : theory and research (77.99%)
Cognitive psychology and its implications (77.64%)
A way of being (77.32%)

Curiously, almost all of the lowest scoring books are business or law titles.

More “stuff like this”…

Just a little follow on from the previous blog post…
Spurred on by comments from Lisa, I’m exploring if we can filter the recommendations so that they become more relevant to students in a specific academic school, or even to students on a specific course, and the initial results look fairly promising 🙂
Let’s look at a couple of examples:
International Journal of Sociology and Social Policy (ISSN 0144-333X)
Here are the recommendations based on usage by all users. A quick browse through the items shows a range of subject areas — social exclusion, economics, human resources, etc, and a student would need to sift through to spot the items relevant to their subject area.
Now let’s filter the recommendations so that they’re only based on usage by students in a specific academic school:

Applied Sciences (which includes tourism and hospitality courses)
Business
Education and Professional Development
Human and Health Sciences
Music, Humanities and Media

…hopefully you can spot that the recommendations suddenly jump to becoming much more relevant to courses in that particular school.
Managerial Auditing Journal (ISSN 0268-6902)
Let’s drill a bit deeper this time and look at courses in the Business school:

Without knowing how our course codes are created, you can probably guess that courses starting with “BA…” are mostly accountancy & finance, and that those starting with “BM…” are to do with leadership and management.

“People who looked at this thing, also looked at this stuff…”

We’ve had serendipity suggestions on the OPAC for nearly 7 years now, but they’ve been based entirely around the physical collection in the library.
After Friday’s Skype chat to the SPLURGE Hackfest, I got to thinking about how we can hook the e-stuff into the recommendations, so I’ve spent the weekend gathering data from our library management system, our link resolver and our EZProxy logs to see what happens if they all go into the same melting pot.
It’s a very rough & ready “crappy prototype”, but you can have a play around with it here. If you get an empty page, click on the “pick random item” link until something interesting happens.
At the moment, the recommendations are being built from a database of just over 5 million events (approx 70% of those are item loans and the rest are accesses of online journals). If you take the “Midwifery” journal as a starting point, you’ll get a list of the other books and journals that people have looked at. The algorithm behind it is the same one I’ve discussed previously.
If you hover over a title, you’ll see the usage info breakdown, e.g. “42 / 56” means that 56 different users in total have looked at the recommended item, and 42 of those also looked at the item we’re generating the recommendations for.
I’ve not done any de-duping, so you might get the same journal title being repeated (once for the print ISSN and once for the e-ISSN), and I’ve not included any ebook usage data yet. I’ve also avoided merging the two lists together until I can figure out a suitable way of weighting book loans against online journal usage.
Picking random items, it’s apparent that some courses lean more towards book borrowing (i.e. very few journal recommendations), whilst stundents studying other subjects are heavy online journal users (i.e. very few book recommendations).
So, what do you think — is it useful to be able to show more than just book recommendations to students?

5 years of book loans and grades (revisited)

Nigel’s comment on the “5 years of book loans and grades” post reminded me that I did do a breakdown by discipline of the same data.
One of the caveats with this is that it represents nearly a decade’s worth of usage and, during that time, the seven academic schools at Huddersfield have changed — e.g. some courses and subjects have moved from one school to another.
Music, Humanities and Media

In terms of books, the students in this school rack up highest average number of loans.
Business

Business students’ borrowing is much lower, but considerably more stable across the 5 years of graduates.
Computing and Engineering

I guess there are no surprises here — when I did my HND in Computing at Huddersfield in the 1990s, I only visited the library once 😀
Education and Professional Development

There’s something interesting going on here — the borrowing levels for firsts and thirds is very similar, with 2:1 and 2:2 being lower. Very curious!
Human & Health Sciences

Applied Sciences

From memory, Applied Sciences make much higher usage of journals than books.
Art, Design and Architecture

The art stock is much more likely to be used within the library, rather than loaned.

CILIP Cymru Conference

I’m journeying down to Llandrindod Wells tomorrow to give a presentation about usage data to the Welsh Libraries, Archives and Museums Conference (hashtag #cilipw11). I’ve been promised that there’ll be real ale there 🙂
You can grab a draft copy of my presentation (“If you want to get laid, go to college…”) from here (15MB).
The main web links in the presentation are:
– JISC Library Impact Data Project
– JISC Activity Data Programme (including a list of the projects)
– Rufus Pollock (Open Data and Componentization, XTech 2007)
– Paul Walk (“The coolest thing to do with your data will be thought of by someone else”)
– University of Huddersfield – Open Data Release (from Dec 2008)

5 years of book loans and grades

I’m just starting to pull our data out for the JISC Library Impact Data Project and I thought it might be interesting to look at 5 years of grades and book loans. Unfortunately, our e-resource usage data and our library visits data only goes back as far as 2005, but our book loan data goes back to the mid 1990s, so we can look at a full 3 years of loans for each graduating students.
The following graph shows the average number of books borrowed by undergrad students who graduated with an specific honour (1, 2:1, 2:2 or 3) in that particular academic year…

…and, to try and tease out any trends, here’s a line graph version….

Just a couple of general comments:

the usage & grade correlation (see original blog post) for books seems to be fairly consistent over the last 5 years, although there is a widening in the usage by the lowest & highest grades
the usage by 2:2 and 3 students seems to be in gradual decline, whilst usage by those who gain the highest grade (1) seems to on the increase

Sliding down the long tail

At a recent event in Edinburgh, I was asked about how we generate the “people who borrowed this, also borrowed…” suggestions in our OPAC and whether or not there are privacy issues with generating them.
Last week, I popped over to Manchester for a meeting of the JISC funded SALT (Surfacing the Academic Long Tail), which is one of the recently funded Activity Data projects. Part of the discussion at the meeting was around how to generate recommendations for items that haven’t circulated many times.
At both events, I promised to put together a blog post detailing the method we use, so here it is!
To generate recommendations for book A, we find every person who’s borrowed that book. Just to simply things, let’s say only 4 people have borrowed that book. We then find every book that those 4 people have borrowed. As a Venn diagram, where each set represents the books borrowed by that person, it’d look like…

To generate useful and relevant recommendations (and also to help protect privacy), we set a threshold and ignore anything below that. So, if we decide to set the threshold at 3 or more, we can ignore anything in the red and orange segments, and just concentrate on the yellow and green intersections…

There’ll always be at least one book in the green intersection — the book we’re generating the recommendations for, so we can ignore that.
If we sort the books that appear in those intersections by how many borrowers they have in common (in descending order), we should get a useful list of recommendations. For example, if we do this for “Social determinants of health (ISBN 9780198565895), we get the following titles (the figures in square brackets is the number of people who borrowed both books and the total number of loans for the suggested book)…

Health promotion: foundations for practice [43 / 1312]
The helping relationship: process and skills [41 / 248]
Skilled interpersonal communication: research, theory and practice [31 / 438]
Public health and health promotion: developing practice [29 / 317]
The sociology of health and illness [29 / 188]
Promoting health: a practical guide [28 / 704]
Sociology: themes and perspectives [28 / 612]
Understanding social problems: issues in social policy [28 / 300]
Psychology: the science of mind and behaviour [27 / 364]
Health policy for health care professionals [25 / 375]

When we trialled generating suggestions this way, we found a couple of issues:

more often than not, the suggested books tend to be ones that are popular and circulate well already — is there a danger that this creates a closed loop, where more relevant but less popular don’t get recommended?
the suggested books are often more general — e.g. the suggestions for a book on MySQL might be ones that cover databases in general, rather than specifically just MySQL

To try and address those concerns, we tweaked the sorting to take into account the total number of times the suggested book has been borrowed. So, if 10 people have borrowed book A and book B, and book B has only been borrowed by 12 people in total, we could imply that there’s a strong link between both books.
If we divide the number of common borrowers (10) with the total number of people who’ve borrowed the suggested book (12), we’ll end up with a figure between 0 and 1 that we can use to sort the titles. Here’s a list that uses 15 and above as the threshold…

Status syndrome : how your social standing directly affects your health [15 / 33]
Values for care practice [15 / 61]
The study of social problems: seven perspectives [18 / 90]
Essentials of human anatomy & physiology [15 / 81]
Human health and disease [15 / 88]
Social problems: an introduction to critical constructionism [15 / 90]
Thinking about social problems: an introduction to constructionist perspectives [21 / 127]
The helping relationship: process and skills [41 / 248]
Health inequality: an introduction to theories, concepts and methods [20 / 122]
The sociology of health and illness [29 / 188]

…and if we used a lower threshold of 5, we’d get…

If you think of the 3 sets of suggestions in terms of the Long Tail, the first set favours popular items that will mostly appear in the green (“head”) section, the second will be further along the tail, and the third, even further along.

As we move along the tail, we begin to favour books that haven’t been borrowed as often and we also begin to see a few more eclectic suggestions appearing (e.g. the “How effective have National Healthy School Standards…” literature based study).
One final factor that we include in our OPAC suggestions is whether or not the suggested book belongs to the same stock collection in the library — if it does, then the book gets a slight boost.

JISC Activity Data Programme

I’m chuffed to bits that the Library Impact Data bid that Huddersfield submitted, along with 7 project partner institutions, was one of the successful ones in the JISC Activity Data Programme and the project will kick off on Tuesday this week!

… the aim of this project is to prove a statistically significant correlation between library usage and student attainment. The project will collect anonymised data from University of Bradford, De Montfort University, University of Exeter, University of Lincoln, Liverpool John Moores University, University of Salford, Teesside University as well as Huddersfield. By identifying subject areas or courses which exhibit low usage of library resources, service improvements can be targeted. Those subject areas or courses which exhibit high usage of library resources can be used as models of good practice.

If you’re interested, keep an eye on the project blog: https://library.hud.ac.uk/blogs/projects/lidp/

Course level journal article feed

Following on from the last blog post, I’ve done some coding to see how well (or not!) a course level new journal article feed might work.
The process behind the code is…

for a given course, identify the most frequently accessed journal titles
use JournalTOCs to fetch the latest articles from each journal’s RSS feed
for each course, create a list of articles (sorted in descending date)

…and you can see the initial output from the code here: http://www.daveyp.com/files/stuff/journals/
For some courses (e.g. Educational Administration) it looks like usage is focused on a single journal, but most seem to bring in content from multiple titles — for example, BSc Criminology is bringing in content from:

The Howard Journal of Criminal Justice (0265-5527)
Nature (0028-0836)
Terrorism and Political Violence (0954-6553)
Criminal Behaviour and Mental Health (0957-9664)
Journal of Investigative Psychology and Offender Profiling (1544-4759)
Crime Prevention and Community Safety (1460-3780)
Criminology (0011-1384)
Crime & Delinquency (0011-1287)
The British Journal of Criminology (0007-0955)
The Police Journal (0032-258X)

One of the opportunities here is to use the journal usage data to identify potentially relavant journals that aren’t being used on a course and include those in the feed. In the above example, the Journal of criminal justice might be such a journal.