I’m taking the Foundations of Data Curation class at GSLIS and just finished a progress report for the MODIS Snow Frequency data set. Here is a link to the report.
https://www.youtube.com/watch?v=QO_AeAX1v5Q
I use data and technology to tell stories that facilitate personal and communal transformation.
I’m taking the Foundations of Data Curation class at GSLIS and just finished a progress report for the MODIS Snow Frequency data set. Here is a link to the report.
https://www.youtube.com/watch?v=QO_AeAX1v5Q
I’ve been working on hard on the Northfield Historical Society website recently and am excited about the progress there! Here’s what I’ve tackled recently.
Next steps:
I’m back for the second day of my internship with the Center for Railroad Photography and Art and wanted to record some of the technical parts of our process.

While most of this collection has already been scanned, some needs to be scanned for the first time. Here is the process I’m using for the scanning process.
I’m scanning using Lake Forest College’s “Epson Perfection V750 Pro” scanner with the frame for negatives. I am scanning at 600 dpi in 16-bit grayscale and saving as jpg. The purpose of these scans is to create decent access copies for the Center with the option to add them to online databases (Flickr or the Center’s website) in the future so that this level of scanning is appropriate.
Most of the files are named using the following convention: Collection Name/Box/Envelope. For example
I’m proposing that we standardize this practice across the entire collection. Additionally – to account for instances where one envelope contains multiple photos, I’m proposing that we extend this convention to be: Collection Name/Box/Envelope/Image. For example,
I think it makes sense to supply the alphabetical distinction only when needed and to use letters instead of numbers because (1) it will improve computer sorting and (2) the original order (within the individual envelopes) is difficult to preserve.
This is the core metadata for the collection
As an aside, I thought it would be fun to mention two pieces of train-specific metadata I’ve encountered so far. I’ve mentioned one already – “MP 60.” Again, I think that means “mile post” but I’m really not sure. I think it would be great to record this and could be very useful to certain people in a certain context. However, it doesn’t fit with other standard vocabularies (city/state) and requires additional context to be useful.
The second train-specific information I encountered is this (2-8-4). A quick google search took me to Wikipedia where I learned that is the Whyte Notation for a particular wheel arrangement (http://en.wikipedia.org/wiki/2-8-4).
I’ve volunteered to speak at the ALCTS Symposium on January 30th about North Park’s experience using demand driven acquisitions for ebooks. The topic of the day is “Collection Directions: The Evolution of Library Collections and Collecting” (website) and is based on the following article:
Dempsey, Lorcan, Constance Malpas, and Brian Lavoie. 2014. “Collection Directions: The Evolution of Library Collections and Collecting” portal: Libraries and the Academy 14,3 (July): 393-423. Links to full text here: http://oclc.org/research/news/2014/10-14.html
The instructions for the day invited us to create a short presentation that “should be practical in nature (what are you doing in your local and/or consortial environment that can serve as a model for other institutions) but also touch on what you think this means (if anything) for the future of collecting.” Because this is my first time ever speaking at such an event, I think I will stay as close to the prompt as possible.
My name is Andy Meyer and I am the Digital Information Specialist at North Park University here in Chicago. North Park University is a small, private liberal arts college with a FTE of less than 2,500. We have a print collection of about 250,000 books that is focused on supporting undergraduate research and graduate study in a few areas. We have a full time staff of nine people. So our context is quite different than some of my fellow panelist. I’m here to reflect on how changes in library collections and collecting are influencing small schools in unique ways.
This is also my first time speaking at such a large forum and – to be honest – I’m a little nervous. So I’m going to stay close to my notes and answer as directly as possible the invitation that I received in December. So be prepared for 15 minutes on the projects I’ve worked in for my institution that can serve as a model for other institutions while touching on what I think that means for the future of collecting.
In particular, I’m going to talk about two different PDA programs that I managed. First, I’m going to talk about a program I lead that helped my library’s reference collection transition from print to electronic. Second, I’m going to talk about the opportunities and challenges of managing a multivendor DDA program, particularly how we planned, implemented, and maintain that program. Lastly, I’ll touch on the next steps we have planned for North Park as well as some broader thoughts about the future of library collections and collecting.
I implemented North Park’s first demand driven program with the head of collection development as a way to help us move our reference collection from print to online. We wanted to provide a lot of eReference content to serve our undergraduate student population and we had to do so with relatively limited funds. This project was relatively straightforward and involved the following steps:
Here is my first bit of “practical advice”: become an expert on your DDA program. We focused on marketing this resource to outside audience (students) but I didn’t anticipate how much I would need to reach out to my co-workers. I learned that I needed to explain everything well and in multiple formats (email, meetings, informal conversation) so that everyone was on the same page. I would imagine this is true and every scale and it was certainly true in our context.
So that is a quick walk through of that process as well as a piece of practical advice. On a broader scale, I think this project raised interesting questions about the use of data in libraries and reduced transactional costs.
An interesting part of this program was a paradoxical lack of data in certain areas and an excess of data in others. Perhaps a better characterization was that we had usage data in a variety of formats that took effort to use and interpret. For example, although we wanted to leverage usage data about our print collection to inform and prioritize our digital purchases, items in our reference collection didn’t have circulation or browsing information. Instead, we had to rely on different sources of usage data – recollections, impressions, stories, etc.
This wealth of anecdotal information and lack of hard data is quite different from the information provided by the eReference platform. Obviously, all you get from them is hard data without any real sense of how they are using it. This data is wonderful but is much more complicated to interpret and evaluate. I’ll return to this point later on.
I thought the easiest way to talk about transactional cost and infrastrucutre was to simply say acquiring thousands of eReference titles was fundamentally different than acquiring thousands of print reference books. This is especially true in terms of cataloging, processing, and physical infrastructure. Using the language of the article, some of this is the difference between print books and ebooks (no physical processing or shifting!). However, the more fundamental shift was not in print vs. electronic but in our reliance on the vendor’s digital network. We relied on their cataloging services. We spot checked a few records and ran reports to get an overall impression, but these process was not were near as complete as our cataloging procedures for print books. However, our evaluation process before loading and on on-going use have confirmed that these vendor supplied records are totally adequate for our use. I think many small schools fear the “flood” of vendor supplied records because it is viewed as an attack on “traditional cataloging” – I’m here to say that in my experience, these fears are largely unfounded: we’ve been pleased with the quality of these records. Furthermore, it would have been wildly impractical to maintain old practices – in fact I would argue that most DDA programs would require a change in processing and (for better or worst) an increased reliances on outside digital networks.
Our major DDA program has a been a multivendor program done for general use ebooks. Our goal in this project was to create a “critical mass” of ebooks so that ebooks would become normal for our student community. We thought that a DDA program could quickly provide this “critical mass” of ebooks as well as be an interesting experiment for us. I want to focus on the word “experiment” because that was and is how I view this DDA program. I also think this language resonated across the library – as a way to demonstrate innovation cost saving to administrators and as a way to assuage certain fears.
Hopefully this overview of the process has implicitly conveyed a point – setting up a DDA program takes a far amount of work at the local level as well as lots of coordination with larger networks.
As a way to pivot from the practical to the philosophical, I want to share a success story from this project. After doing all this work, I was naturally very curious to see what our first DDA loan would be. Our first loan of this program was “Religious Ethics and Migration: Doing Justice to Undocumented Workers.” This was a tremendous success for a number of reasons:
So our first DDA loan was a perfect fit for our collection. These felt like great news (and a relief!) and showed how this new service could extend our collection to meet the needs of our community.
So our first loan was a tremendous success…but it was also a little unsettling and raised questions: “Why didn’t we order an electronic copy of this book (it’s a North Park author and a required text)? Do we want a print copy and an electronic copy? Also – I wasn’t aware that this record was even in our DDA pool. It just shows that the lines between the library, vendor, and our community were blurred. This “blurring” is particularly clear when I process the file of new records, delete the old records, and process the short term loans and purchases. It’s very apparent that we don’t have local control of this collection in the same way we control our print collection.
Overall, this project is 100% doable for small institutions with limited budgets and staff time. I’ve provided a brief overview or our experience so far – feel free to reach out to me directly with additional questions or concerns – I’m happy to share more. As I said, this program is still on-going and I’d like to conclude my time with two questions that I’ll be exploring when this experiment is over.
In both of these projects, we’ve assumed that print collection and circulation information would correspond to eReference and eBook usage. This assumption seems logically but also seems somewhat troubling. By looking at the data and talking to my community, I’d like to see how closely print usage matches electronic usage. While I anticipate general correspondence, I also wouldn’t be surprised to see striking differences – differences that perhaps to the changing patterns of research and learning mentioned in this article and elsewhere.
I’ve talked a bit about how success in our PDA programs relied heavily on partnerships between our institution and the different vendors. In particular, I’ve mentioned what is essentially outsourced cataloging, shared technical infrastructure, and the blurred lines between library collection and vendor service.
However, North Park also exists in a highly networked environment with our consortial network – CARLI – and we haven’t fully studying the impacts of our DDA programs on our consortial network. How do the various DDA programs of member institutions affect the consortium as a whole – especially in terms of a shared catalog and interlibrary loan? Essentially, I’m curious to know how
Overall, for a small liberal arts school, these experiments in DDA have been very positive. We’ve been able to provide a lot of content to our community at relatively little cost, we have learned a lot from this experiment, and we are thinking more critically about the future of collections and collection development.
I’m excited to start an internship working with the Center for Railroad Photography & Art and the Lake Forest College Archives! I want to use this space to write up a bit about the Center for Railroad Photography & Art, the collection I’ll be working with, and the some initial thoughts and goals for my particular project.
The Center for Railroad Photography & Art is a organization dedicated to the preservation and interpretation of railroad art as it intersects with American life. There website provides much greater detail about their mission: http://www.railphoto-art.org/about/. They are involved in publication – including the journal “Railroad Heritage” – and maintain an online web portal that serves as a digital collection of railroad photography.
I’ll be working directly with one of the Center’s collections – The Fred M. Springer Collection. The Center for Railroad photography & Art has compiled a relatively complete biographic section – and he seems like a pretty interesting guy! I’m particularly fascinated by his interest in narrow gauge trains. I’m working with his collection of black and white negatives that are currently housed and curated by the Lake Forest College Archives.
The collection is roughly 7,500 Cellulose Acetate Film Base negatives – most seem to be “Kodak Safety” negatives. Most of the negatives are 3.5 inches wide and 2.25 inches tall. These negatives are currently housed in sleeves and envelopes of various sizes and materials (paper, plastic, etc.) and I do not know if they are archival quality of not. These sleeves contain critical metadata for the collection. The collection seems in very good shape – no signs of deterioration at this point.
Options/Recommendations
The physical negatives are currently arranged alphabetically by railroad (and possibly by train number following that) in what is presumed to be the Springer’s original organization. The collection is also physically divided into 6 boxes. Digital files are currently named used a locally devised scheme of Collection_Box#_Envelope# and seems to correspond closely to the existing physical arrangement.
Options/Recommendations
Implement metadata standards that relate to the intended purpose and audience of this collection. Existing metadata is written on the envelope and often includes: train, train number, location, date, and sometimes a short description.
Options/Recommendations
To organize my thoughts and get a better sense of the project I wanted to compile a list of questions for the Center for Railroad Photography & Art. Here is a quick list.
I’m new to the area and wanted to begin building a body of knowledge and resources as it relates to this project. Here are some helpful resources I’ve found so far.
I just watched R. David Lankes’s recent keynote and wanted to post a short review of the material he covered.
“Publisher of the Community: We’re All Doomed” Closing keynote for the NISO Workshop on “Using the Web as an E-Content Distribution Platform: Challenges and Opportunities.” http://quartz.syr.edu/blog/?p=6265
Abstract: We need to build platforms for scholarship and knowledge development, not information and content delivery. These platforms are not about APIs and eContent, but about people and content. We need to strive not for discovery, but epiphanies.
[vimeo id=”109836897″]
I really enjoyed this presentation and will take a few new ideas from it. Beyond new ideas, it was refreshing and inspiring and perhaps a bit challenging to hear his thoughts about collections and technology.
Lankes began by talking about the “death of documents” in that the definitions of documents are changing in rapid and new ways. Digital stuff is different at a fundamental level and the editorial process is radically different because digital information resources are living, changing things. There is continuity in this change but, overall, the change has been a massive paradigm shift. The words “ebooks” and “ejournals” are really just metaphors – we are trying to fit new information into old information formats.
This section flooded my mind with thoughts.
Lankes also argues for a “sea change” in the history of library and library services – we need to stop thinking about content. There have been massive shifts in the last 40 years:
A Netflix for education doesn’t work because education isn’t based on consumption – it’s based on learning and learning is based on conversation and participation.
However, I do take issue with some of these claims. First, this presentation created vocational confusion. Lankes argues – persuasively, I might add – that libraries must move beyond a focus on collections to a focus on service and even beyond that to community conversations. However, my library and my position in the library are organized around the first two areas – collections and services – and my job is focused squarely on maintaining access to our collections and a few of our services. So while I’m inspired by the claims that we need to focus beyond those things I wondered what those new services and platforms for community conversations could actually look like and – as importantly – how do we move in that direction without the required skills? Lankes said that “People create restrictions that technology has removed” but that gives me pause for several reasons. First – the technical challenges to what Lankes proposed are not trivial. The technical infrastructure takes time and energy and expertise that many libraries and community groups simply lack. Second, how do we manage this change?
These – of course – are criticisms as much as they are struggling to think through the ramifications of his argument in my local context and in the context of my future.

Looking forward, I’d like to think more about the following things and perhaps explore them in library school:
Okay. I think that’s enough in the way of reflections. That’s plenty to think about and plenty to wrestle with for now.
Brian C. Gray presenting Analyzing and Selecting the Best Discovery Solution for Your Users and Organization(s). Library website: http://library.case.edu/ksl/.
How to make a decision? Gray worked with OHIOLink and started with a list of specifications as well as how all the user audiences would work with each service. Wanted the search to be comprehensive, not federated, able to be embedded in other library resources, and ease of maintenance. Gray worked with a small group of people and focused on making decisions quickly.
What are the challenges? All the products are very different. Very different. Librarians and average library user have very different expectations and viewpoints. Hard to answer the question – “what is this tool searching?” – because of the complexities in the index. Thinks about the diversity of users – can you provide customized tools for different communities? Discovery process changes rapidly – perhaps too fast for librarians to manage. How difficult is it to change and maintain our holdings? Have someone that works with all users, someone with cataloging/metadata expertise, IT and webmaster, financial agent.
Advice: Take advantage of trials. Make sure staffing is sufficient. Don’t take too long in making the decision. Do local inventories: processes, technology, expertise, time, and financial resources. Define your local goals! Gray’s goals included: drive usage to certain resources, increase resource usage, change user behaviors, help people brainstorm ideas, speed up research, challenge the “google mindset,” change library instruction, reduce the number of access points, provide a common tool, what’s your long term plan, what’s your level of commitment? Beyond goals, prioritize things. Things like: user needs, local customizations, addition costs, what level of control do you have, multilingual interface, user customization options, how does it work with backend management tools, usage statistics, facets customization, what level of support do you require, can you add local content?
Create a list of specifications; list them and then score them (yes/no or a numeric scale) and what are the “must have” features. Gray has a collection of great rubrics for evaluating discovery services.
Don’t take forever to implement – start in beta and make changes accordingly. You won’t be perfect, it will get better. Work with the vendors and take full advantage of their skills. Expect constant change.
Dana Belcher from East Central University presenting on “How Discovery Services Tie into University Missions.”
Why discovery? Lots of reasons! Foster a learning environment, educate students, keep up with changes in society. These resources are tied to the University’s missions statement. Moved the library webpages totally to LibGuides; a positive change. Selected EBSCO discovery because they already had a lot of EBSCO products and the price is right. Make sure you understand 100% of the products/services offered! They created a profile for each major/area to customize a database list. After that, they worked with EBSCO to better define the limiters and facets – especially making changes to better meet user needs. Instruction is focused on limiting/faceting a search instead of “choose this database” type instruction.
Advice:
Focus on meeting the needs of the University and the needs of students. Everything is tied to the strategic plan for the library.
I watched this webinar as part of the I LEAD USA program sponsered by the Illinois State Library Association. Here is a quick review to document the experience and reflect a little bit.
The presentation addressed the question “How do we create a culture of experimentation?” and started with a thought experience involve a student interview set in the year 2018 that praises a program call “Library that learns you.” This program – sort of like an embedded librarian/google information/netflix recommendation engine is praised. And then the speaker asked two questions: “Is this a good idea” and “What do we need to learn/develop for this to work”
[youtube id=”rtWrnuHt9RQ”]
I think this exercise encouraged people to think creatively about long term goals. There was some limited discussion about the limits of library services (especially related to privacy and big data) and whether this “service” is an extension of traditional library services like embedded librarianship. People also talked about the need for cross-campus connection and the technology learning curve.
The presenter then asked simply “what is experimentation?” and gave some examples for science and engineering. In general, these examples stressed the need for trial and error and the combination of serendipity and hard work. He then offered these principles.
He then framed the rest of the presentation in three parts (I’ve used them as heading below)
2. Practical examples from libraries. (physical space and learning space)
3. Things we can do to create a better culture of experimenting
Next steps:
Last week, I watched the webinar entitled “What is a data-driven academic library?” hosted by Library Journal and really, really enjoyed it. Like, best webinar ever.
Sarah Tudesco (@studesco) and Bonnie Tijerina (@bonlth) did a wonderful job on the #ljdatadriven webinar! #BestWebinarEver?
— Andy Meyer (@ameyer24) December 4, 2013
I was too engrossed in the content (and busy tweeting the event) to take comprehensive notes. But here are some of my notes from the webinar as well as later reflections.
To be data driven means to use data to drive change.
I feel like this is the core of the presentation. To be data driven means to use the data libraries and library systems collect to drive decision making. She proposed a five part method to accomplish this goal.
1. Create Questions
All questions are good – but bigger questions might need to be broken down into smaller chucks. Good questions are ones that are answerable based on data. To me, that meant questions like “are the nursing databases meeting current needs” must be translated into questions about usage and coverage – and that we should be honest and clear about that work. She also advised to not be overly reductive and to pay attention to other factors. Lastly, she advised people to let these questions guide your research – don’t get lost in the data!
2. Create a Plan
Along with crafting questions, Tudesco suggesting creating a plan. Essentially, what data will answer your question and how will you get that data? Refine your questions and create a plan or timetable for this project.
3. Collect and Manage Data
Honestly, I feel like I usually begin my data-driven projects at this step – and my projects almost always suffer because of this! My workflow usually begins with “What can I do with this data?” rather than the more important “How can this data help me make a decision?”. Allowing the question and project drive the data needs – and not the other way around – is a very valuable lesson I will take from this webinar.
She was also clear that sometimes the data you want/need doesn’t exist and that perhaps you need to create a tool to gather that data. Whether than means a survey or structured observations will be guided by the question
4. Analyze the Data
Again, this methodology focused on using the data to make a decision. This data analysis will bring together a lot of information – perhaps from different sources and with different nuances – and this will require interpretation and analysis to make sense of this. More on this later.
5. Make a Decision
This is where all that data matters! It’s very basic but it’s worth repeating – Data Driven libraries use data to drive decisions.
Tudesco mentioned three core skills library staff need to truly embrace the data movement and mentioned several tools under each area. This section also provided my favorite quote – one near and dear to my heart!
#ljdatadriven "Excel is my best friend"
— Andy Meyer (@ameyer24) December 4, 2013
I’m trying to remember what she presented the best I can while also taking the liberty to add my own skills and tools to this list. This is a work in progress and will likely add more in the future.
Data Capture
Data Analysis
Presentation
Becoming a great storyteller is just as important as accurate and pretty charts for libraries leveraging data — @studesco #ljdatadriven
— ER&L (@ERandL) December 4, 2013
Overall, this was a wonderful and timely webinar. I will do my best to take the lessons and tips to heart and structure all future data project in light of what this presentation taught me.
Here is the link to the webinar – I believe that if you register you will get archived access to webinar so you can watch it yourself!