Showing posts with label duplicate. Show all posts
Showing posts with label duplicate. Show all posts

Saturday, March 3, 2012

Mitigating a Data Verification Nightmare: Thoughts on Removing Duplicate Files

I had gotten myself into a data nightmare, with a bunch of files that appeared partly duplicative of one another.  I wanted to get rid of the duplicates.  This appeared likely to be a long struggle.  This post is one battle in that war.

For starters, I used DoubleKiller (I had the pro version, but the freeware one would have helped too) to pluck the low-hanging fruit -- to delete, that is, the verifiably exact duplicates.  But now there were files with almost identical names (e.g., Longfilename and LongfilenameA), files with identical names but different extensions (e.g., was Filename.pdf simply a PDF version of Filename.doc?), filenames with slight differences (e.g., was 2010-09-10 Résumé an essentially identical duplicate of 2010-09-10 Resume if their times were identical but their sizes were different?), and so forth.

It was easy enough to just guess at it and delete the ones that looked like they might be duplicates.  In many cases, that would have been fine; it wouldn't have made any real-world difference.  Obviously, though, this would not be not a good data management solution.  More like data abdication.  Second-best, I could identify likely duplicates and put them in a ZIP file, out of the way.  One problem with ZIP files, I had found, was that the reason for zipping them could fade from recollection, over a period of years; and then, one bright day, someone might decide to see what was in there, and the monster would live again.

One general underlying purpose was to have files that would actually be useful.  I thought that the processes of gradually absorbing useful files and eliminating unnecessary ones might be aided if I could sort them by topic.  Here, again, there were some obvious ways of quickly taking care of large numbers of files, such as those that were already sorted into folders with meaningful names.  But that left quite a few that were not usefully categorized.

In some cases, I could categorize files just from the information in their names.  But I hated to spend the time to do it manually.  I tried to sort them into categories by identifying key multiword phrases they contained.  That problem became complicated by variations in punctuation and other textual vagaries.  Therefore, I started over, this time beginning with an effort to clean up punctuation and other aspects of the text.

While that cleanup attempt was underway, I also looked for ways to reduce the number of filenames being sorted.  This brought me back to a focus on identifying duplicates.  It seemed, belatedly, that it would have been useful to have named all files according to a consistent rubric.  I took a look at that in a separate post.  That got me to a point where I was able to name many of the files in a certain standard way.  So, for example, an email would be named using a Date-From-To-Subject format.

Along the way, certain realizations forced themselves into my consciousness.  One was that, as a general rule, if it looks like a mess, it is probably a mess in ways that you haven't even imagined.  A corollary is that whatever you do, you will have to do over again, once you have discovered additional unforeseen ways in which the data are intertwined.  I did find that doing it wrong, several times over, was a good way to become familiarized with possible starting points.  Having a good backup was essential.  Being able to reconstruct steps, by saving relevant documents in generational steps, was a real plus.

There was also the problem of deciding whether to nibble around the edges or make a decisive stroke to divide major problem areas from one another.  The quandary here was that the decisive stroke would likely go astray if it was not informed by prior familiarity with the actual fault lines shooting through the data.  In the worst case, you would not only add to the confusion, but would divide things in exactly the wrong direction, so that the procedures taken with Group A would have to be repeated with Group B -- except that, inevitably, they could not be repeated *exactly* with Group B, because the two groups would differ in some subtle but significant way.  But nibbling around the edges could turn relatively simple aspects of the project into enormously tedious exercises in repetition, as exceedingly minor pieces of the puzzle were interminably quasi-resolved.

I made these notes during a particularly dark moment in the process.  I am pleased to report that, later on, when I returned to these notes to wrap them up and post them, I had turned to other projects, and was thus making good progress toward resolving the data verification nightmare by simply ignoring it until it (or I) went away.

Friday, March 18, 2011

Thunderbird for Windows: Transition from Portable to Desktop; Duplicate Email Remover

I was using Thunderbird Portable 3.1.4 in Windows 7.  I wanted to use an add-on (Remove Duplicate Messages (Alternate) 0.3.6) to delete duplicate email messages.  I got the impression that it wouldn't run on the portable version.  I had been thinking about switching to the desktop version of Thunderbird anyway, and now seemed like the time.  To figure out how to transition from portable to installed versions of Thunderbird, I ran a search and found advice that seemed on point.  I did not precisely track all of the steps I took in this process, but the following is a pretty close approximation.

I started by installing regular (i.e., not portable) Thunderbird.  I think I created an email account at that point.  This generated C:\Users\Administrator\AppData\Roaming\Thunderbird\Profiles\f0xqaflh.default.  (The f0xqaflh part was randomly generated -- other installations would have a different ????????.default file.)  I closed Thunderbird and moved C:\Users\Administrator\AppData\Roaming\Thunderbird\Profiles\f0xqaflh.default to D:\Thunderbird\Profiles\f0xqaflh.default.  I put it on D so that it would be saved in case of Windows reinstallation.

Then I went to Start > Run > "thunderbird.exe -ProfileManager."  In Profile Manager, I clicked on Create Profile > Next > Choose Folder and pointed to D:\Thunderbird\Profiles.  I exited Profile Manager and moved the contents of ThunderbirdPortable\Data\profile (i.e., just the profile subfolder) to D:\Thunderbird\Profiles.  I clicked on my Start Menu shortcut for Thunderbird (not portable).  It ran, and it seemed that all of my emails were there.  I deleted the folder containing the portable version.

I hoped this was all I needed.  Now it was time to try to delete duplicate emails.  I installed the duplicate email remover add-on (Tools > Add-ons > Extensions tab > Install) and ran it (Tools > Remove Duplicates).  It wouldn't check my archive folder until I turned off the Skip Special Folders option (Tools > Add-ons > Extensions tab > Options > Message Comparison tab).  At first, I used the default comparison criteria in that same tab:  Author, Recipients, CC List, Message ID, Send Time, Size, Body, and Subject.  This did not identify too many duplicates, but it appeared they were exact duplicates, so I could delete them all without much manual comparison.  I ran another search, without the Message ID criterion, and yet another, without the Size comparison.  The former likewise seemed not to require much manual comparison; the latter did.  In other words, the final comparison criteria (Author, Recipients, CC List, Send Time, Subject) produced many alleged duplicates, some of which were of very different size.

The add-on did not allow me to open individual emails (via double-click or right-click), to see why two emails bearing the same subject, date, time, etc. would be so radically different in size, so I had to do a lot of manual toggling back and forth between the duplicate remover and Thunderbird, and then searching for individual items in T-bird, to check emails one by one.  In this regard, it was not like DoubleKiller, which I had found to be an excellent duplicate file finder.  But the manual selection process was similar:  check or uncheck the desired item under the "Keep?" column.  Both of these programs would probably have been easier to use if it had been possible to select or deselect items by clicking anywhere on the line, rather than having to mouse over to precisely the checkbox spot each time.

The add-on did allow arrow-key and spacebar navigation and selection.  Playing with this, I eventually discovered that the Enter key would open T-bird to one of the identified duplicate messages, but in that case the comparison window disappeared and I was back in Thunderbird, leaving me to wonder why I was now seeing only one of the duplicates.  Then I realized, oops, hitting the spacebar had not actually opened the selected duplicate; it had gone ahead and run the deletion.  Well, I hoped those 700 messages really were duplicates.  I had been verging toward just saying to hell with the time-consuming and awkward manual comparison process anyway; I just wasn't quite ready for this to happen.  I looked in Thunderbird's Trash folder and realized that I had not emptied the trash before running the duplicate checker (another ideal feature for the duplicate checker), so now I would have to restore not just the 700 messages that I had apparently just deleted, without an "Are you sure?" message, but would also have to restore about 700 other messages that were apparently in the Trash previously, since I was now seeing a total of 1400 messages there.  As I looked at the Trash, I found myself wondering, actually, what was wrong with those 700 other messages.  They didn't seem to be messages that I would have wanted to delete, unless they too were duplicates.  I decided to move the whole lot of them to the archive folder that I had been dup-checking.  At this point, needless to say, I was beginning to fear that I might just be turning my whole email archive into a giant hash.  I started back through a sequence of dup-checks, beginning with the most conservative (i.e., with the most comparison criteria checked), but of course this time I had no patience for checking individual items.  Instead, I just dreamt of an update that would actually display large thumbnails of alleged duplicates, right there in the add-on.

The column headings in the dup-check results window permitted sorting in ascending or descending order.  At first, I thought that feature was not working for some criteria.  Then I figured out that it was meant to sort only within a comparison.  For example, if Size was not a comparison criterion, it would not be in boldface in the top row, and then clicking on it would sort alleged duplicates according to size; but if Size was a comparison criterion, it would be bolded, and then clicking on that heading in the top row would do nothing, since in that case all duplicates within a set would be identical by definition.  It would have been helpful if selected comparison criteria headings had enabled a sorting of all pairs.  That is, if I was comparing by Send Time, I wanted to be able to show the earliest ones (i.e., the pairs of allegedly time-identical messages) first, so that I wouldn't have to do so much jumping-around when I toggled to Thunderbird for a manual comparison.

After running the several comparisons mentioned above, I tried running one with only the Send Time and Subject criteria checked.  This revealed some apparent duplicates whose only difference was that for some reason one item in a pair would be enclosed in quotation marks (e.g., a message from "Joe") while the other would not (e.g., a message from Joe).

That was the end of my use of the add-on at this point.  I returned to finish this post several hours after completing these processes.  It appeared, at that point, that the transition to desktop Thunderbird and the use of the add-on to delete duplicate emails were both successful.

Friday, January 21, 2011

Windows 7: The INSTALL Partition

For a long time, probably since the 1980s, I had kept my Windows installation on drive C and my data on drive D.  (Drives, or partitions of drives, can be readily created within Windows 7 and also by partition manager programs.  GParted, included in the downloadable Ubuntu CD, has been the most reliable partitioning programs for my purposes in recent years.)

Having the data on drive D had several advantages.  One was that I could back up drives C and D on different schedules.  Once I got a good Windows installation set up on drive C, I didn't need to back it up very often.  I'd just make a drive image (using Acronis True Image in the past few years), and I'd make an updated image backup just when I had done some significant new program installation or adjustment.  By contrast, I would want to back up my data on drive D on at least a daily basis.  Of course, when I made the image of drive C, I would need someplace to put it other than drive C.  An external drive was a possibility, assuming the bootable CD would recognize it, but it would tend to be slower, and this would be complete downtime, when it would not be possible to do other work.  It worked the other way, too:  having the data on a separate partition meant that I could completely reinstall Windows without affecting or even worrying about my data.

There were exceptions to that last statement.  Some program configurations were so detailed, time-consuming, and/or oft-changing as to constitute a sort of data.  Not the kind of data I was supposed to be working on, but data nonetheless.  An example:  Firefox add-ons.  It could take a half-hour or more to find, install, and configure my Firefox add-ons after doing a new Windows installation.  Some add-ons allowed me to export my saved settings, but obviously I would not want to store those on drive C; they'd be wiped out if I reinstalled Windows sometime down the line.

I also found it was handy to keep a local copy of the programs that I would install in Windows, after installing the Windows operating system itself.  I did not want to have to re-download all those programs and re-invent all of the things I had previously figured out about installing them.  So instead I had folders containing the programs to install, with installation notes in accompanying text files, and I named the folders in such a way as to guide me in the installation sequence that worked best (e.g., "01 Motherboard Drivers").  There were also quite a few program installers that I tried and uninstalled, or hadn't gotten around to installing.  Also, some ISOs -- ready-to-burn CD images that I had downloaded but hadn't burned to CD, or wanted to keep because it was a hassle to re-download a 700MB image.  I had collected these sorts of things, not only for Windows, but also for Ubuntu.

I accumulated about 75GB of this stuff.  Keeping it all on drive D meant that my daily data backups were swelling up with all this material that didn't need to be backed up every day.  So at some point I moved a bunch of it over to its own partition, with occasional backups.  What I kept on drive D was mostly stuff that I was actually using in my current installation.  One example:  those Firefox settings files.

Sometime in the late 1990s, I discovered that I could move my Start Menu to drive D.  This would have the advantages mentioned above, including especially the fact that my custom-arranged Start Menu (top-level folders:  Productivity, Online, Multimedia, Tools, Startup, Miscellany) would not have to be rearranged each time I installed Windows.  As long as I installed everything in its default installation location, the shortcuts in the Start Menu would come back to life as soon as I reinstalled the target program where the shortcut expected to find it.

A few months before writing this post, I came to realize that, of course, I could also install my portable applications in the Start Menu.  This would be ungainly in the sense that a bottom-level folder in the Start Menu might contain a slew of program files instead of a nice, orderly collection of shortcuts.  But it was handy for keeping everything that ran in one place, where I could copy it to a jump drive and use considerable parts of it on any other Windows machine.  The Start Menu, wherever located, could also be accessed and synchronized on a network, so that I only needed to configure the Start Menu once and would then have it available for any computer I would attach to my home network.  (Making it available did require a registry tweak.)

Putting the portable applications in the Start Menu had an unwanted side effect.  Many portable apps (especially those coordinated by PortableApps.com) used many of the same program files.  That was a problem because I liked to use DoubleKiller to delete duplicate files on drive D.  I couldn't do that anymore, at this point, because it would detect tons of duplicative portable program files that were supposed to be there.  I looked for a different duplicate remover, one that would allow me to exclude folders like the Start Menu folder, but ultimately decided to stay with DoubleKiller for now.  The reason was that I did not want to risk that, one fine day, I would forget to exclude the Start Menu from a DoubleKiller sweep, and (although this was unlikely) would punch the wrong button and delete that Start Menu from my hard drive.

What I decided to do, instead, was to move the Start Menu so that it would join those other program files -- programs to be installed, etc. -- in their own partition.  I called it INSTALL and gave it a letter of W, so that its location would not be affected by the connecting and disconnecting of various USB drives and whatnot.  So then hopefully the shortcuts in it would stay in place, and would require no further adjustment forever and ever.

Right now, unfortunately, they did require adjustment.  I modified the registry tweak to point toward drive W:\Start Menu rather than D:\Installation\Start Menu.  This caused almost no problems.  The main issue was just that a bunch of shortcuts were now dysfunctional, in that they pointed at executable files on D that were now on W.  I ran Glary Registry Repair 3.3.  It identified 165 new registry errors.  As I scrolled down the list, I noticed that a huge number of the problems identified by Glary were links to IrfanView on D.  I considered doing a global registry search and replace for the IrfanView location -- or, indeed, for all references, changing them from D to W, perhaps with the aid of a global registry search-and-replace tool like Registry Toolkit ($25) or Registry Replacer ($15) or Replace Registry Values (free).  But then I decided it might be safer to let Glary fix those dud links and then do a search in a registry editor for any remaining references to D:\Installation|Start Menu.  I used O&O RegEditor to do that search.  I was now down to a total of just 29 registry references to D:\Installation\Start Menu.  I was going to edit them in O&O, but then decided not to.  As long as they weren't hurting me, I was better off just leaving them alone.  One ill-advised registry edit could cost me an hour or more for recover or restoration.

So this was pretty much the end of the project, aside from some continuing cleanup, correction of shortcuts, etc.  I now had an INSTALL partition, labeled as drive W, containing my customized, shared Start Menu and installers for various programs.  I could now run DoubleKiller on drive D without worrying that it might knock out program files, and without having to do manual exclusions to focus it on the data duplicates that I was seeking.

Thursday, January 20, 2011

Windows 7: Choosing a Duplicate File Finder

I had long used DoubleKiller to find and delete duplicate files.  The switch to Windows 7 made this seem like a good time to review the market and see if there were better alternatives.

This question acquired urgency because I came across a need to reconcile two hard drives, and discovered that I could not.  The reason was that my customized Start Menu -- containing not only links to programs but also the entire program folders for my portable apps -- was rife with duplicate files.  These were not like duplicates among my data files.  There, I would generally want to delete duplicates.  Here, deleting duplicates would mean that programs would not run.

DoubleKiller permitted me to compare entire drives and to mark tons of duplicates for deletion, all in one move.  That sort of thing could obviously be misused.  But I had gotten reasonably good at focusing it so as to delete only what I wanted to delete.  What I wanted now -- what DoubleKiller did not offer -- was a way to exclude some subfolders within those drives and top-level folders.

This was a delicate matter.  After all, in some operations (e.g., reconciling two hard drives), I would be trusting this program to accurately identify and delete thousands of duplicate files with one click.  Judging from the number of freeware duplicate detection utilities, many of which drew approval from multiple reviewers, this was a kind of task that could be done well by a good programmer.  I just didn't want any unexpected surprises.  Fortunately, at this point I was using Beyond Compare to check my drives against backups for changes on a file-by-file basis, and I had also begun using GoodSync to analyze differences between two computers (more specifically, to stop before proceeding if more than 10% of the files differed), so I had some protections.  Nonetheless, I wanted a reliable program.

So I went looking for a replacement for DoubleKiller -- something that would render it duplicative, if you will.  Ranked by frequency of download, SnapFiles said the top five general-purpose (i.e., not image-specific) duplicate detectors were (Auslogics) Duplicate File Finder, Fast Duplicate File Finder, AllDup, Duplicate Cleaner, and LookDisk.  (DoubleKiller was sixth.)  eHow gave instructions for using Duplicate Cleaner, Auslogics Duplicate File Finder, and AntiTwin.  CNET was a muddled source for this purpose, offering general-purpose utilities (e.g., Glary) with some duplicate detection capabilities; I decided to stick with a utility dedicated to this specific task.  In two different searches of CNET, Duplicate Cleaner and Auslogics Duplicate File Finder seemed to be leaders.  A review at Gizmo named Duplicate Cleaner, Auslogics, Anti-Twin, and Fast Duplicate File Finder as the top four.  SnapFiles had not said how many downloads they had; I decided their selection was iffy.

I went to the homepages for Duplicate Cleaner, Auslogics, AntiTwin, and Fast Duplicate File Finder, looking particularly for flexibility in excluding the files within a subfolder from deletion.  From the webpage for Duplicate Cleaner, it sounded like they had this capabilty; from those for Auslogics and Fast Duplicate File Finder, I could not tell.  Anti-Twin sounded rather bare-bones -- saying, for instance, that "the software ignores file names."  For some projects, I wanted the option of comparing file names, to identify nominal duplicates that I would then manually choose among.

Conveniently, then, Duplicate Cleaner seemed to be not only the leading program named by several of the sources cited above, but also the only one among these contenders that said, right on the wrapper, that it had what I wanted.  I liked that they had a manual on their website, and that they also had the beginnings of a support forum.  I downloaded and installed it.  Its interface was more user-friendly, which meant that I immediately disliked it.  Seriously, I was turned off by the options to compare for "Same Artist," "Same Album," etc.  But I figured I would learn to ignore that sort of thing.  Its user friendliness did make it much easier to select drives or folders for comparison; and when I selected a drive and then tried to select one of that drive's folders, the program asked if I was naming the folder for exclusion.  Exactly what I wanted.  In the "More Options" area, the default comparison was MD5, but it had an option for byte-to-byte and two SHA formats.

I ran a check for duplicates, using both Duplicate Cleaner (version 2.0) and also DoubleKiller (version 1.6) on the same drives.  I had Duplicate Cleaner set to search for Same Content, any date, no file filters.  I wasn't sure how to set the file size criterion so I set it to Any Size.  It gave an option of setting a minimum file size of 1KB, but I didn't want it to skip files that were, say, 500 bytes; I had written many short text files, containing a note on this or that, that would be that size or smaller.  I went with the default MD5 content comparison type.  In DoubleKiller, I searched for files with identical sizes and CRC32 checksums, and I told it to ignore system files (because I didn't want a hundred copies of Thumbs.db or desktop.ini) as well as files whose size was equal to 0 KB.  I also told DoubleKiller to exclude .dll, .sys, vxd, and .inf files, as well as those of zero bytes and system files.  In this area, I felt that DoubleKiller's options were better, though not ideal.

I started both searches on the same machine at the same time.  During the searches, DoubleKiller remained visible, and slowly started adding visible duplicates to its list.  In a previous search, Duplicate Cleaner had seemed to stall.  Its title bar said "Not responding" and it got that sort of half-screwed-up look that graphics would get sometimes, when the computer was running out of memory or about to crash.  This time around, that didn't happen; still, it was hard to rouse it from the taskbar.  Duplicate Cleaner did not show any interesting specifics about duplicates while it was running.  Both used a screwy way of calculating the percentage of completion:  wait forever for the job to get to 1% completed, then soon it's at 5% and by the time it gets to 40% it's moving right along.  Perhaps because CRC32 checksums were easier to calculate, DoubleKiller was done first, scanning 198,645 files in 46 minutes.  Duplicate Cleaner carried on for a total of 57 minutes.  Both programs surely would have done the job much faster if they had not been competing against each other for the same computer resources.  Duplicate Cleaner found "56 Groups of duplicates" involving a total of 160 files.  DoubleKiller did not present a number, but by my count, it found 111 duplicative files.

I found the screen font displaying outputs to be much more readable in Duplicate Cleaner.  Duplicate Cleaner also allowed me to change the background colors behind alternating duplicate pairs, so as to make it easier to figure out where the alleged duplication was appearing.  Unfortunately, Duplicate Cleaner had many columns of information, but right-clicking on the column row did nothing.  In other words, it did not appear possible to suppress those columns so as to focus on the ones that mattered (e.g., MD5 hash) without scrolling right and left for every duplicative pair.

I was not satisfied with Duplicate Cleaner, and I was out of time to compare the others.  I decided, for now, to continue with DoubleKiller, and to review one of the others sometime in the future.