Showing posts with label folksonomies. Show all posts
Showing posts with label folksonomies. Show all posts

Friday, May 23, 2008

Assault and Ambiguity

It began with a chance encounter. I was walking through a room with a TV on and a news caption was running during a commercial break. The newscaster intoned with provocative seriousness how a woman had followed a man who had sexually assaulted her weeks earlier to his home and he had been arrested. The location was my town and it rang a bell.

Three weeks back my wife and I went to a doctor’s appointment in Lafayette and, as we returned and entered our neighborhood, we were surprised to see a half-dozen police cars in our otherwise perfectly tame patchwork of planned homes and parks. During a walk two hours later there was still a squad car on one side street. A scan of the police blotter turned up the cause: a woman walking with her toddler had been assaulted and groped by a man who ran away when struck in the groin with a sippy cup.

Weeks passed until I overheard the news of the arrest. Good for her! The next phase of datamining impressed me with the thoroughness of the picture. I was able to use the television station archives combined with Google to find the mug shots, the original sketch of the suspect done by police sketch artists, the suspect’s arrest status in the county courts system, the location of the alleged perpetrator’s house, the suspect’s father’s name and place of business, the suspect’s mother’s name and place of business, a previous citation of the suspect for a moving violation (infraction) in a neighboring city, the county records concerning the amount and type of mortgage held on the suspect’s home, and a satellite view of the home as well.

Amusingly, also, was that the reporter in the news piece actually drove by our house and coincidently filmed our various vehicles. I could likely have read the license plates if I wanted from the footage.

Overall, I had managed to scour out all the corners of ambiguity concerning when, where, who and how, leaving only the strange question of why left in fuzziness. Why was this 24-year-old still living at home, jogging at midday and preying on middle-aged women? Why was he living in this neighborhood where even a Megan’s Law offender is fairly hard to find?

But strangely, it was the suspect’s last name that was the key to developing the search picture because the last name was so unusual. Had he been “Jim Smith” or “Joe Sanchez” or “Mike White” it would have been virtually impossible to make as much headway in extinguishing ambiguity. George Miller is quoted something like “There is only one problem in Artificial Intelligence: words have more than one meaning.” (And I can’t resolve the ambiguity of the source of that quote because George Miller is too ambiguous). This problem is amplified for searching across identities of places and people, or when special identifiers are introduced as placeholders in a single document (this happens quite often in technical literature where an acronym is used locally as a technical shorthand but is ambiguous outside that document or domain). Moving to the level of folksonomies for, say, labeling pictures on the web, we see the problem exasperated by the natural telegraphic shorthand that any labeling scheme suggests to the user purely by dint of the size of the entry fields.

Clever approaches to trying to apply context to help address these limitations start with statistical co-occurrence-based disambiguation and linkage analysis, and then run all the way through to using complex ontologies to try to infer the best relabeling of the ambiguous entity or concept as a canonical identifier. None of these methods can hope to achieve any level of perfection but a basket of them can enhance the process of information discovery and disambiguation.

Thursday, April 24, 2008

Folksontamasticons and Ambiguity

Folks might not be all bad, though. For instance, in my Ofamind technology, this blog and social bookmarking sites like del.icio.us, the tags that are attached to documents serve to help people find and retrieve information. Tagging is a counterpoint to the idea of structured ontologies and metadata because it builds from the ground up rather than from the top down. The term coined for these tagging schemes is “folksonomy.”

But are folksonomies useful and consistent? Some studies suggest they are useful under some circumstances. For instance, querying across the titles and descriptions using tag keywords on del.icio.us bookmarks results in a precision-recall of only 50%. In other words, the tags are not also in the texts around 50% of the time, and so provide an additional channel of information for retrieval. People appear to think differently about tags than they do about titles and descriptions.

In terms of consistency, however, a very large number of tags are used only once or are used in differing and inconsistent ways that indicate ambiguity over multiple user subcommunities. Examples might be “architecture” used to refer to computer architecture and building design, or “camp” referring to drama or outdoor recreation.

A couple of interesting questions emerge about how to refine the power of folksonomies.

For instance, can the title and description (or full blog content) be used to automatically suggest tags that are based on other tagging schemes? The Ofamind system partially does this by automatically categorizing web content among your “views” or collections as you surf. It does a fair job, too, for a great deal of content. This can be seen as a personalized metadata tagging filter, since the view association to content is essentially a categorical tag.

Similarly, business taxonomies, controlled vocabularies, full ontologies and other mechanisms could be used at authoring time to try to suggest or overlay more consistent tags onto web content, enhancing searchability and even supporting reasoning about content. For Ofamind, a subproblem that we are currently working on is how to disambiguate extracted people, places and organizations in order to produce high-quality metadata using a combination of human tagging and automatic methods.

Then the folksonomy becomes more of a folksontamasticon, combining folksonomy, ontology and onamasticon in a rare new tag.

Wednesday, November 14, 2007

Folksonomics and Conceptual Metadata

I necessarily think about information design as a part of my professional duties. I also try to keep frosty on novel ways that information might be presented. So I naturally have been revisiting these notions in some avocational research I have been delving into concerning health care reform.

Health care reform is, of course, a deeply political topic that breaks down along several ideological and interest dimensions. In supporting the claims for all sides, basic research is mined and often cherry-picked to build a case. And my point in this entry is not to make an ideological or political claim but to describe in a way how that information is discovered, used and reused.

Now I was quite a novice on the topic of health care economics and reform ideas when I began researching the topic. I certainly had personal negative and positive experiences over the years. I also had heard the reviews and blowback over Michael Moore’s Sicko (though have not yet seen the film). But, beyond that, I had no real understanding of who the players were, what the research suggested, or what the counterclaims were.

My understanding built from a range of sources, most of which were simply not accessible even a decade ago to casual researchers like me. Instead, you had to be a Beltway insider who subscribed to think tank newsletters and research publications. But now I can download and read Commonwealth Fund reports, CATO news briefs, and a host of other resources and become a moderately well-informed amateur researcher. I can access huge swaths of blogs and commentary, reflecting different perspectives. I can even organize the information by collecting it together and then labeling it for easy recovery based on a recollection.

What I can’t easily do yet is to be able to answer specific questions that have not already been answered in some publication, but that emerge out of the collected information. For example, after I read how the German medical system did not have the kinds of rationing and waits for access that we associate with certain aspects of the British and Canadian systems in a Commonwealth Fund report, I wanted to know the details of the German system. Was it a single-payer or nationalized health service? Perhaps a hybrid? What was the role of doctors and information technology? It took quite a lot of searching to finally be able to answer those questions, ultimately using a Siemens Medical Technology prĂ©cis and market analysis of the German system.

Could approaches like structured metadata via Semantic Web technologies assist me in these tasks? Perhaps, but it seems to require that propositional information above the level of named entity extraction could be accurately indexed. For the first question, I would need documents labeled with “Structure of the German Medical System” or the equivalent rather than the many nuanced and varied ways that we write. Moreover, the proposition needs to express the timeliness of the resource, a problem I frequently encounter when trying to fix problems with my Linux computers. How timely is a given piece of information?

We can’t, however, expect individuals to be able to code that metadata in any consistent way, though I believe there is a folksonomic method that can lend a hand: a community of users can gradually improve the metadata structure and content in much the same way that Wikipedia is gradually improved. The Wikipedia model has also shown that quality can be maintained—with fits and starts—by a community of users and some policies in place.