Sunday, February 9, 2014

Metadata: Good, Bad, or all in how you use it?

The subject of metadata has definitely been making its rounds in the news of late, often in a more negative than positive light. Between the collection of personal metadata by governmental agencies, and the theft of metadata that led to the compromising of millions of Target customers' credit card information, we are all more aware of how vital metadata can be to our privacy and security, and the potentially malicious ways in which it can be gathered and utilized. But are these instances enough to cause a shift in the way our metadata is treated, and overshadow the great number of uses for it? Two of my LS 566 classmates have posted recent blogs that have well represented the extremes of the Good/Bad scale that metadata application seems to slide along.

Heather Castle's blog post, Metadata and the Olympics provides links and discussion about metadata collection occurring during the 2014 Sochi Winter Games. Among the less-than-surprising revelations is the disclosure of significant amounts of metadata being collected by the Russian government concerning personal information for users of communication and information networks at the Olympics. Such data includes personally identifiable information, and even payment for products purchased and services rendered. Ostensibly being utilized to ensure security of the Olympic Games, a potentially high-profile target for nefarious or malicious activities, this disclosure still represents the less exciting uses for metadata that we are seeing more and more of on a regular basis.

On the other end of the spectrum, a recent post by Danielle Tinkler discusses the use of metadata in tracking flu activity across the globe. Illustrating information such as the severity (quantity) and dispersion of flu cases, both nationally and internationally, this use of metadata has tremendous potential value in keeping the public aware of the potential for infection, but also in assisting health officials to decide how best to utilize the resources at their disposal.

As with any tool, the threat, or value, of metadata has far more to do with the motivations of the people collecting and presenting it, than it does the data itself. If we seek to curtail the collection of metadata that potentially poses a threat to us, do we run the risk of equally preventing the use of metadata that provides value? Can we have the latter without exposing ourselves to the former?

Friday, February 7, 2014

Metadata Concepts: Dublin Core


One of the important concepts mentioned frequently in our current text reading is Dublin Core. Dublin Core (DC) is a metadata initiative and element set resulting from a 1995 metadata workshop, in Dublin, Ohio, conducted by the OCLC and NCSA. It has been adopted as an ANSI/NISO, ISO, and IETF standard. DC was originally conceived and created as an effort to extend a standard metadata set to the description of electronic documents, as standards of the time skewed heavily toward the description of physical items. Rather than as a replacement to existing standards, Dublin Core was designed more as a complementary element set, capable of “filling in the gaps” for document and object types that other metadata standards failed to adequately cover. Compatible with a variety of coding languages, including HTML and XML, DC produces simple, structured records which can be used independently, or in conjunction with other metadata standards (i.e. MARC). The simplicity of using the element set, coupled with its ability to be modified, extended, and interoperably used with other standards, has led to its widespread acceptance and use in the metadata community.

The base Dublin Core set consists of fifteen elements:
            -Title
                -Creator
                -Subject
                -Description
                -Publisher
                -Contributor
                -Date
                -Type
                -Format
                -Identifier
                -Source
                -Language
                -Relation
                -Coverage
                -Rights

These core elements are all optional to use, and can also be repeated by the metadata creator. The focus of the elements trends toward the areas of object search/retrieval, and resource discovery, but the extendable nature of Dublin Core enables the set to be adapted and modified to the needs of specific domains and communities. Sets of extra elements, known as qualifiers, can be utilized in conjunction with the base elements to extend the scope of the set in a standardized way. Examples include the Canberra Qualifiers: Language, Scheme, and Type, which expand upon the functionality of the base elements, but cannot be used independently.

Dublin Core plays a significant role in the world of metadata standards, especially as applied to electronic documents and media. Its simplicity and adaptability allow it to be tailored to the individual needs of specific information communities, as is illustrated the set's mapping to schema such as MODS and MADS, among others. The set continues to undergo development and revision through the fostering of the Dublin Core Metadata Initiative (DCMI).

Monday, February 3, 2014

Misusing metadata to get the most out of your classical music

Metadata isn't something sacred, although I sometimes can't shake the feeling that using it incorrectly is violating some kind of unwritten cosmic law. It's a tool, and like most tools, you want to use it in the way that it was intended, or your results may skew from the expected. Musical metadata is no exception. Song titles, album titles, artists, composers, and genres are all informational tools that, when used properly, should organize your collection, collocate (or relate together) similar items, and make searching for, and locating, specific items an efficient and painless process.

But sometimes, your music software or device just doesn't work the way you want it to, and the only real solution is changing the metadata from something that's technically correct to something that works a little better in practice. This blog entry will touch on some of the reasons I've had to "fix" my metadata in ways that might not quite be right, but have provided me with the functionality I desired. Be warned though, brave souls, the Naxos Blog discourages you from going to too great a length in disturbing your musical metadata.


1. Problems with device or software functionality
Sometimes, you need to break your metadata because your device or software just won't work when it's entered correctly. Having recently purchased a Sansa Zip Clip mp3 player in my effort to throw off the Apple yoke (ok, I still patronize iTunes), I was unpleasantly surprised to learn that my new gem of a music player doesn't recognize the "Disc #" metadata field for a multi-disc set. By not recognizing this tag, my player could not understand that the six track 1s for the set actually belonged to six different discs, and instead treated them all as tracks from the same disc, completely compromising the player's ability to play tracks in the correct order. While perhaps less of a problem for rock or pop albums, for a set of symphonies, this was devastating.

To fix this, I was left with two alternatives: break the multi-disc set up so that each disc was treated as a separate album (i.e. Beethoven: The Symphonies [Disc 1], Beethoven: The Symphonies [Disc 2], etc...), or alter the track numbering scheme. The former, while perhaps doing less to compromise the integrity of the metadata, would have severe implications on my ability to quickly search or sort through albums (imagine adding an extra 30-40 album titles to your list), so I chose the latter. Employing a numbering scheme by which all tracks in the set were consecutively, and uniquely, numbered (i.e. disc 6 has tracks #35-43 rather than #1-8), I was able to restore the compilation to regain its proper ordering for both my iTunes-based devices, as well as my Zip Clip.

2. Customizing naming conventions
Consistency of metadata is vital to its effective use, both in enabling informational items to be grouped and sorted, but also searched for. When dealing with the description of classical works, naming conventions for an item can often vary wildly. Consider the final movement of Brahms' Third Symphony, titled, as most movements are, for its tempo, "Allegro." The number of classical movements entitled "Allegro" surely counts in the thousands, if not more, making the task of proper identification and retrieval of this particular piece extremely problematic if that is the only information used for the title. But what information should then be included in the title to ease the process? Potential options include:

Symphony No. 3-Allegro
Third Symphony-Allegro
Brahms' Third Symphony-Allegro
Brahms: Third Symphony-Allegro
Brahms: Symphony No. 3 in F major-Allegro
Brahms: Symphony No. 3 in F major, Op.90 - 4. Allegro

All are technically correct, and vary in description from relatively simple to rather complex. For my purposes, I used the last of the naming conventions listed, as I trend toward detailed descriptions. The format does have some drawbacks, however. The use of the composer name at the beginning of the description helps with sorting, but should only technically appear in the "Composer" field. Also of concern is the number of characters used, which could exceed the displayed character limit on certain devices, and prevent the full title from being seen. Additionally, ensuring that all pieces in the collection conform to this format is an extremely time consuming process, and with so much information being entered, attention to detail is vital to ensure each title is entered correctly.

This may be less a misuse of metadata then simply trying to get the most out of it, but custom naming conventions often require a significant amount of metadata to be altered. Whatever you may decide to use, try to pick a format early in the process, and stick to it.

3. Using incorrect fields to ease searching/sorting
One trick I've seen mentioned on many classical music blogs and forums is the deliberate use of metadata in improper fields to ease the searching/sorting process. For much modern music, albums will have a single artist, while possibly many songwriters or composers, making "Artist" a preferred sorting/search field. For classical music, this is reversed, with "Composer" being a relatively singular field for an album, while having many potential artists. This can leave the classical listener with a relatively more difficult search when using software and devices that place greater emphasis on the "Artist" field, especially if their collection contains both classical and modern music. To fix this problem, some classical collectors will swap the "Composer" and "Artist" metadata entries, greatly easing the browsing process.

These are a few of the "wrong" uses of metadata I've come across that can actually help the functionality of your devices or collections. While there are definitely some drawbacks to employing them, whether it be the time involved, or the difficulties caused by updates to software or devices, there are also some tangible benefits. I should also point out that if you upload your metadata to be used by other people, you should strongly consider NOT entering information incorrectly in your collection, so as to preserve the integrity of the metadata going out across the web. I would be greatly interested to hear of any fixes that any of you may have employed in your own music collections, or perhaps in other uses of metadata.

Sunday, February 2, 2014

Dvořák or Dvorak? Does it really matter?

One of the more significant issues I've run into in bringing some order to my classical music collection is the lack of consistency in metadata entry. Whether it be the composer, artist, or even name of the piece, standards for how metadata is entered seem to vary wildly from CD to CD, wreaking a fair amount of havoc on my ability to organize, sort, and search my collection. But this problem with metadata goes beyond just the description of items in music software, rather extending into the description of information items across the internet, and the whole digital environment as well.

To apply the title question as an example, let's look at the Czech composer Antonín Dvořák. You will probably notice the name is one that is not incredibly friendly to our American keyboards, with accents over the "i" in Antonín and "a" in Dvořák as well as the hacek, or hook, over the "r" of the latter. It only makes sense, then, despite altering the integrity of the name, that we will often prefer to enter it here as the anglicized Antonin Dvorak instead. Regardless of which form of the name is more convenient, however, if you're striving for consistency in data entry, what form should we be looking to use in the "Composer" field of a program like iTunes, Media Monkey, or Music Collector? A browse through the information that my collection of CDs actually produced belies an even bigger problem than expected:

Antonín Dvořák
Antonin Dvorak
Dvořák
Dvorak
Dvořák, Antonín 
Dvorak, Antonin
Dvořák, Antonín (1841-1904)

One can imagine how much more difficult such a trend has on my ability to quickly and efficiently find the particular works I'm looking for when browsing by Composer. Now multiply this by the 20-30 classical composers that are included in the collection, and the sometimes-odd use of the Composer field in modern music, and you've got a true mess on your hands, and a tool that is almost unusable. This doesn't even consider the similar problems that one might encounter in the Title or Artist fields. 

The fix for such a problem, in the library world, is known as authority control. Authority control is the reason that you can utilize a library catalog, and efficiently find results on a particular topic through the use of subject and author search terms that have been pre-defined by organizations such as the Library of Congress in their LC Subject Headings and LC Name Authority File. For this particular case, the LC Name Authority File uses the official entry: Dvořák, Antonín, 1841-1904, which ensures that any works created by Dvořák are uniformly described and easily searchable.

Authority control is an easy fix, in theory, but it also requires the establishment of a person or group to create and maintain the authority files, and the willingness of the affected metadata creators to accept the use of authority records. While it is possible that companies of a like field (i.e. music companies) could agree to the common use of authority records, attempting to similarly gain the allegiance of the millions of us common folk that create metadata on a daily basis is a much more severe task. I don't see a simple solution to this problem, especially considering the multi-national character of the internet and its associated metadata. It may just be that the online environment is too far gone to hope for any meaningful return to organization and control, if there ever was any. What do you think? Do you have any authority control horror stories? 

Friday, January 31, 2014

Metadata and Classical Music

As an aficionado, and burgeoning collector, of classical music, the topic of metadata in digital music collections has become a sort of pet project of mine. Having recently gone through the process of re-importing my entire classical collection (a small one, admittedly, around 1000 songs), I can attest to the fact that attempting to adequately manage the metadata of such a collection is a labor of love. Inconsistent data entry, missing or incorrect metadata, and some unique problems that hamper classical listeners in comparison to those with mainly rock/pop collections, all conspired against my efforts to bring some much needed order and cohesion to my music files.

In a recent blog, LS 566 classmate Mary Elizabeth Watson touched on the subject of metadata and digital classical music collections, and, with reference to the Naxos Blog on classical music, pointed to some of the more significant hurdles that classical listeners must face in organizing their digital collections. Some of the problematic areas mentioned include: the lack of authority control across metadata entries, the incompatibility of metadata tags and their functionality across varying music software and devices, and the purposeful misuse of metadata in order to work around those functionality issues.

One of the fundamental difficulties classical listeners face when dealing with their digital collections is that, by and large, digital music platforms, devices, and software were made with the classical genre as a bit of an afterthought, if thought of at all. Software such as iTunes emphasizes Artists, Albums, and especially single Songs as primary categories, which have limited, or differing, usefulness to a classical listener, who may prefer organizing and searching collections by Composer or complete Work/Piece (i.e. Beethoven's Symphony #5 rather than just the 2nd Movement). Audio quality, perhaps not a primary consideration when listening to the latest Katy Perry single, but paramount when trying to squeeze every ounce out of Mahler, is often compromised by the lack of support for certain lossless file formats. In other instances, the incompatibility of certain metadata tags (i.e. the Disk # tag in iTunes) across players/devices makes attaining even the proper playback order for a piece impossible. While there are a lot of software options out there, all of them seem to have their drawbacks, whether it be: overly simplistic, overly complex, poor compatibility with your purchased music device, lack of compatibility with downloaded music file types, or a myriad of other reasons.

While the picture may look slightly bleak for classical listeners, there's actually great reason for optimism, however. There has never seemed to be a better selection of recordings available, at a better price, thanks in part to digital distribution. Listeners have unprecedented control over the organization and description of their personal collections, depending on how much time they want to sink in to the process. And on the software front, software platforms have gradually begun to include more and more features that are beneficial to classical collectors in particular (i.e. iTunes improved functionality for multi-disc sets, options for gapless playback, lossless audio compression, etc...). But there is still a good deal of work to be done (and a true software platform/device built for the classical genre wouldn't be bad either), and one of the areas that can definitely be improved is musical metadata. Next blog will take a look at one of those problem areas...authority control.


Tuesday, January 28, 2014

Metadata Opportunities for Librarians

As I inch closer toward completion of my MLIS, thoughts invariably turn to what comes next in my career. Though I have a great deal of interest in a traditional library or archival job, including areas such as reference or cataloging, I am also increasingly aware of the possibilities of library and information professionals outside the traditional walls of our field. In a recent blog post, LS 566 classmate Molly Porter discussed her work as an archivist for the NASA's Marshall Space Flight Center, and some of the non-traditional and metadata-related activities that increasingly characterize her work. Examples include: the "crafting of a metadata strategy for the center's film and media migration project," continued development and updating of web content, as well as increased social-media activities of varying kinds.

Having had the fantastic opportunity to experience a similar work environment during an internship with NASA's Jet Propulsion Laboratory, I can echo the increasing role that metadata plays in the world of information professionals, and the excellent chance it affords us to branch out from our traditional role inside the library. The tasks I was responsible for were two-fold: metadata description of digitized documents into a topical digital archive, and the digitization of project documents and reports, and their subsequent entry into the project's document repository. The former job had me manipulating metadata in a manner that was distinctly similar to the controlled process of library cataloging, including the application of document titles, authors, and even subject headings (created from a archive-specific taxonomy). The latter project, by contrast, bore more similarities to records management and preservation, and any controls placed on the metadata entered were a result of my own preferences. This task, especially, illustrated to me just how important consistency of metadata entry can be to the organization of a database or records repository, and how much responsibility the metadata creator bears for it.

It is exciting that library and information professionals are increasingly afforded the opportunity of metadata-related jobs, not just with scientific organizations such as NASA, but also in the corporate world, with health and medical organizations, and countless other environments in which the manipulation and organization of metadata, records, and digital documents require our background. Would love to hear from any of you that regularly work with metadata, or have been able to find your librarian skills in demand outside the library walls.

Sunday, January 26, 2014

Metadata Concepts: MARC

Over the course of my LS 566 blogging, I hope to be able to touch on some basic metadata concepts, schemas, standards, and other tools that are important to the library and metadata fields. For the first of these such posts, I'd like to focus on an integral component of metadata in a library setting- MARC.

MARC, short for MAchine-Readable Cataloging is a data format standard first developed during the 1960s, and still used today in its MARC 21 form. MARC provided information professionals with a format which enabled the creation, use, and transmission of electronic catalog records for libraries, and thus made the use of computerized library catalogs, and the sharing of catalog records, viable. As both a national and international standard for bibliographic data, MARC has had a tremendous on the cataloging world, and is virtually synonymous with the concept of library catalog records.

In form, MARC is a series of 3-digit, numerical fields, or tags, with each such tag representing a piece of bibliographic information, such as the: title, author, or subjects for a given work. Tags are modified through the use of indicators and subfields, which allow the entry of additional information, or denote that the record be used or read in a certain way (i.e. indicators can be used to prevent a record from being searchable in the library catalog, while another tells the computer to skip a certain number of characters in the entry, etc...).

A very simple example of a single MARC tag:

245 14 $aThe Three Musketeers / $cby Alexandre Dumas ; with an introduction by Allan Massie

In this example:
-245 is the field for the title of the book
-14 is a series of two indicators: 1 noting that the title should be added to the catalog, the 4 that the first 4 characters of the entry ("The" and the space after) should be skipped (articles like "The" and "A" make organizing and searching for materials are problematic, and so are ignored by the catalog
-$a is the subfield in which the title of the work is entered
-$c is the subfield for the statement of responsibility, which in this case, includes the author and writer of the introduction

Though complicated to use and understand, at times, MARC has had a tremendous impact on the development of modern cataloging. While it has served the library community well for the better part of 40 years, the possibility of change is on the horizon. Among other factors, a new set of cataloging rules (RDA) is steering the cataloging world in a new direction, and it is questionable whether MARC is the best fit going forward. I hope to touch on this subject in a future blog.

For more information about the MARC standard, there are two great resources online: the MARC Standards page, through the Library of Congress, and OCLC's Bibliographic Formats and Standards page.