Ok, I really hope I'm not wasting anyone's time here. This seems like such a basic issue that it's more likely that I'm wrong than the code is. Anyway.
I'm testing with a bare Pressflow 6.16.77 install, with nothing but taxonomy_xml 6.x-2.x-dev installed.
I've also tested against the HEAD of taxonomy_xml, but the import process couldn't complete.
The attached test case file passes the W3C's SKOS Basic integrity test case. The Thesaurus compatibility test case throws warnings about terms having the same prefLabels, but it doesn't error.
It's got four terms in it, three of which have the same prefLabel:
<skos:Concept rdf:nodeID="06420a84bba83106ff1c0088025abf12">
<skos:inScheme rdf:resource="http://skeyn.com/schemes/EDVOC/"/>
<skos:prefLabel>Abbey College</skos:prefLabel>
</skos:Concept>
<skos:Concept rdf:nodeID="b35451a8306cec3017afb9a5da941eef">
<skos:inScheme rdf:resource="http://skeyn.com/schemes/EDVOC/"/>
<skos:prefLabel>Abbey College</skos:prefLabel>
</skos:Concept>
<skos:Concept rdf:nodeID="e7fbc23e2dbdb73325d4c4fb1fd557e6">
<skos:inScheme rdf:resource="http://skeyn.com/schemes/EDVOC/"/>
<skos:prefLabel>Abbey College</skos:prefLabel>
</skos:Concept>
<skos:Concept rdf:nodeID="d5cc68af141a1d041fd7410f0ea02fe1">
<skos:inScheme rdf:resource="http://skeyn.com/schemes/EDVOC/"/>
<skos:prefLabel>Abbey College, Ramsey</skos:prefLabel>
</skos:Concept>
My problem is that once I've imported the data, I only have two terms - "Abbey College" and "Abbey College, Ramsey" - even though the nodeIDs for all four terms are unique. Renaming to "Abbey College1" etc. results in a successful import.
The only reference to non-unique prefLabels I can find is in the W3C's SKOS primer:
Following common practice in KOS design, the preferred label of a concept may also be used to unambiguously represent this concept within a KOS and its applications. So even though the SKOS data model does not formally enforce it, it is recommended that no two concepts in the same KOS be given the same preferred lexical label for any given language tag.
Drupal doesn't require term names to be unique (I can add the terms manually), and as I read this, SKOS doesn't either. So, what am I missing?
| Comment | File | Size | Author |
|---|---|---|---|
| nsSchool_skos_small.xml_.txt | 1.67 KB | dotton |
Comments
Comment #1
dotton commentedI figured it out - there's a checkbox under advanced. I'm just going to slink away now.
Comment #2
dman commentedUse-cases differ on whether a soft text match or a strict ID match are appropriate to use, and (as you've found) what to do in case of conflicts during imports.
There's more to it, but I supported the best case I could, concentrating on better support for re-importing of the same data set over top of itself ... mainly bescause that was the one I was testing a lot.
Note that ideally, I'd not suggest using rdf:nodeID (Which - I think - is local to the parse process) but rdf:about or rdf:id (or is it just 'id'?) ... um , to identify, uniquely, a single concept in a re-import-safe manner. NodeIDs mean nothing to this system, so using them for uniqueness is futile.
Whether this works is another matter, but the dream would be that the ID is saved the first time through, and recognised for an update the second time.
Lacking that, it does a text-only match on the label. And, using that advanced option you found, either create or update.
Comment #3
dman commentedLooking at the spec I understand a little more how rdf:nodeID represents a 'blank node' inside a document.
As such, what the code is doing is correct by not retaining it. a bnode - that is, a data object identified by a rdf:nodeID - is labelled only for syntax structure within a document, and has no relation to concepts outside that document, nor persistence or identity that can be referred to.
bnodes and rdf:NodeID properties are fine for self-contained documents that need to internally say that A links to B and B links to A, but are NOT to be used to say that A has any properties or identity.
... It looks like some FOAF implementations have had the same confusion. As indeed did some folk at W3C - on nodeID
Comment #4
dotton commentedI'm finding that every time I uncover a "bug", it's actually a hole in my/our understanding of how this stuff works. The learning curve is a killer.
That last link is fascinating. Thanks. One question, though.
The "scope" of NodeID is the file. Fair enough. And NodeID is an "anchor" you can use to add structure to the file (relationships between concepts).
But if you can't extract those relationships in some form, what's the point in having them at all? (I should expand, here - when we use NodeID, taxonomy_xml imports the concepts as a straight list - that is, it doesn't import the relationships between them. This is something of a moot point as I've worked around the problem, but I'm still curious).