Problem/Motivation
When a table has exactly 20023 lines, it is very unsettling to read "… imported 19987 …"
Steps to reproduce
This is most likely caused by duplicate "unique" fields in the source, but for the average user it's quite an ordeal to walk some 20k records to see why this actually happened.
Shortcut for Mac/Linux users: from the shell use
cat YOUR_FILE.csv |awk -F";" '{ print <strong>$1</strong> }' |uniq -d
where $1 is the number of the "unique" field (here the first column of the table was considered holding unique values). Just in case it's not so unique, you'll get a list of non-unique content, like so:
107116
112102
116105
118305
118403
120201
121192Proposed resolution
It would be nice if Feeds would log skipped lines like so
Feeds skipped non unique field "FIELD_CONTENT" (line xy of imported_file)
Remaining tasks
Write the code :)
Comments
Comment #2
megachrizFeeds doesn't keep track of unique values
Feeds doesn't keep track of which unique values it has seen during an import. So if you configure a source field as unique and it appears to be not unique, Feeds considers an item with the same value on the source field as the same record. If you have configured to do not update existing items, then Feeds skips it because it sees there's already an entity with that identifier on the system. It doesn't know that entity was imported during the same import. If you have configured to update existing items, then Feeds will overwrite the values on the entity, meaning that the first record's data is lost.
Example:
When the source field "ID" is marked as unique:
Add line xy of imported file to log message
This sounds like a good idea to me: that you can see in the logs the item number. I agree that would be helpful especially in case of CSV imports.
Comment #3
nofue commentedSorry I wasn't all to clear in my description. I do not demand that a CSV importer should care for unique fields, but it would be helpful to know why there is a difference between lines read and lines imported. Even a short message might help, but of course a list of duplicate lines would be the icing of the cake.
Comment #4
megachriz@nofue
No problem, I just tried to explain how things currently work in Feeds. :) This way you know that Feeds detecting duplicate lines may be a lot more work to get implemented and that only adding a line number to a log message looks more doable to achieve.
Comment #5
nofue commentedSo then: Thanks for the crash course -- I know I badly need it :)