<?xml version="1.0" encoding="utf-8"?>
<rss version="2.0">
<channel>
<title>Digital Preservation Q&amp;A - Recent questions tagged disk-image</title>
<link>https://qanda.digipres.org/tag/disk-image</link>
<description>Powered by Question2Answer</description>
<item>
<title>How/where to store metadata about optical media sector layout in METS/PREMIS</title>
<link>https://qanda.digipres.org/1146/where-store-metadata-about-optical-media-sector-layout-premis</link>
<description>&lt;p&gt;
	I'm drafting a METS/PREMIS profile for images/rips of optical media images (ISOs for data sessions; WAV or FLAC files for audio). One of the pieces of metadata I'd like to include is the output of the &lt;a href=&quot;https://linux.die.net/man/1/cd-info&quot; rel=&quot;nofollow&quot;&gt;&lt;span style=&quot;color: #0000ff;&quot;&gt;cd-info tool&lt;/span&gt;&lt;/a&gt;, which contains information about the sector layout of the disc. Here's an example:&lt;br&gt;
	&lt;br&gt;
	&lt;a href=&quot;https://gist.github.com/bitsgalore/9a2838481574c040f7c4b7da4ed59926&quot; rel=&quot;nofollow&quot;&gt;https://gist.github.com/bitsgalore/9a2838481574c040f7c4b7da4ed59926&lt;/a&gt;&lt;br&gt;
	&lt;br&gt;
	However I'm unsure how (and where) to store this info in METS. My initial idea was something like this:&lt;/p&gt;
&lt;ul&gt;
	&lt;li&gt;
		Create a METS &lt;span style=&quot;font-style: italic;&quot;&gt;techMD&lt;/span&gt; element which is associated with the structmap &lt;span style=&quot;font-style: italic;&quot;&gt;div&lt;/span&gt; element that encompasses all files that were extracted from the physical disc (typically one ISO image and/or multiple audio files).&lt;/li&gt;
	&lt;li&gt;
		Inside this &lt;span style=&quot;font-style: italic;&quot;&gt;techMD &lt;/span&gt;element, create a PREMIS&amp;nbsp; &lt;span style=&quot;font-style: italic;&quot;&gt;OBJECT&lt;/span&gt; instance with &lt;span style=&quot;font-family:courier new,courier,monospace;&quot;&gt;&lt;span style=&quot;color: rgb(0, 128, 0);&quot;&gt;xsi:type=&quot;premis:representation&quot;&lt;/span&gt;&lt;/span&gt; (since it describes &lt;span style=&quot;font-weight: bold;&quot;&gt;the disc as a whole&lt;/span&gt;, and not an individual ISO image or audio file!)&amp;nbsp;&lt;/li&gt;
	&lt;li&gt;
		Then use PREMIS unit 1.5.7 &lt;span style=&quot;font-style: italic;&quot;&gt;objectCharacteristicsExtension&lt;/span&gt; as a container for wrapping the iso-info output.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;
	&lt;br&gt;
	The problem here is that the &lt;a href=&quot;https://www.loc.gov/standards/premis/v3/premis-3-0-datadictionary-only.pdf&quot; rel=&quot;nofollow&quot;&gt;&lt;span style=&quot;color: #0000ff;&quot;&gt;PREMIS 3.0 data dictionary&lt;/span&gt;&lt;/a&gt; says that the &lt;span style=&quot;font-style: italic;&quot;&gt;objectCharacteristicsExtension&lt;/span&gt; unit (and also its parent unit 1.5 &lt;span style=&quot;font-style: italic;&quot;&gt;objectCharacteristics&lt;/span&gt;) is &quot;Not Applicable&quot; for the intellectual Entity and Representation object types!&lt;br&gt;
	&lt;br&gt;
	This makes me wonder how others are handling this. Is this an oversight of PREMIS, or is there some other (possibly&amp;nbsp;&lt;br&gt;
	better) way to do this that I've overlooked?&lt;br&gt;
	&lt;br&gt;
	Any suggestions appreciated!&lt;/p&gt;</description>
<guid isPermaLink="true">https://qanda.digipres.org/1146/where-store-metadata-about-optical-media-sector-layout-premis</guid>
<pubDate>Mon, 05 Feb 2018 16:09:33 +0000</pubDate>
</item>
<item>
<title>Parity data for ISO images: anyone doing this? Best practices?</title>
<link>https://qanda.digipres.org/1117/parity-data-for-iso-images-anyone-doing-this-best-practices</link>
<description>&lt;p&gt;
	Earlier today I had a discussion with a colleague on possible ways to store (ISO) images of CD-ROMs and DVDs in our repository system. In addition to checksums, he suggested to also generate and store &lt;a rel=&quot;nofollow&quot; href=&quot;https://en.wikipedia.org/wiki/Parchive&quot;&gt;parity data&lt;/a&gt; for each image file (e.g. using the par2 tool: &lt;a rel=&quot;nofollow&quot; href=&quot;http://manpages.ubuntu.com/manpages/trusty/man1/par2.1.html&quot;&gt;http://manpages.ubuntu.com/manpages/trusty/man1/par2.1.html&lt;/a&gt;). This would enable one to repair files in case of bit-level corruption.&lt;br&gt;
	&lt;br&gt;
	This made me wonder how commonly parity info is used in digital archives, and if there are any best (or at least recommended) practices. E.g. what levels of redundancy are typically used? I'm not really familiar with this at all, and it's not a subject that is often mentioned in digital preservation discussions (but see &lt;a rel=&quot;nofollow&quot; href=&quot;https://twitter.com/anjacks0n/status/733281959762395139&quot;&gt;https://twitter.com/anjacks0n/status/733281959762395139&lt;/a&gt;).&lt;/p&gt;</description>
<guid isPermaLink="true">https://qanda.digipres.org/1117/parity-data-for-iso-images-anyone-doing-this-best-practices</guid>
<pubDate>Wed, 25 May 2016 13:48:28 +0000</pubDate>
</item>
<item>
<title>Incomplete ISO image after imaging CD-ROM - how to prevent and detect this?</title>
<link>https://qanda.digipres.org/1076/incomplete-image-after-imaging-rom-prevent-and-detect-this</link>
<description>&lt;p&gt;
	While running some tests creating CD-ROM ISO images with ddrescue, I ended up with ISO images that were incomplete in some cases (last ~50 MB of image file missing), &lt;em&gt;even though ddrescue’s log file didn’t report any errors&lt;/em&gt;. Below the results I got from 4 attempts at imaging the same CD-ROM on the same PC (note that some of the ddrescue options I used are slightly different, but this appears to be unrelated to my issue). For this I used 2 different external DVD readers:&lt;/p&gt;
&lt;ol&gt;
	&lt;li&gt;
		Reader A - modern Samsung USB device;&lt;/li&gt;
	&lt;li&gt;
		Reader B - old SATA (internal) device, refurbished to USB.&lt;/li&gt;
&lt;/ol&gt;
&lt;h3&gt;
	Attempt 1 - reader A&lt;/h3&gt;
&lt;p&gt;
	Command line:&lt;/p&gt;
&lt;pre&gt;
&lt;code&gt;ddrescue -b 2048 -r4 -v /dev/sr0 windows_98_upgrade_nl.iso windows_98_upgrade_nl.log
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;
	This resulted in a 601.7 MB ISO image. Here are the contents of the log file:&lt;/p&gt;
&lt;pre&gt;
&lt;code&gt;# Rescue Logfile. Created by GNU ddrescue version 1.17
# Command line: ddrescue -b 2048 -r4 -v /dev/sr0 windows_98_upgrade_nl.iso windows_98_upgrade_nl.log
# current_pos  current_status
0x23DC0000     +
#      pos        size  status
0x00000000  0x23DCB000  +
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;
	I.e. the log file indicates the CD was imaged without problems. MD5 checksum is &lt;code&gt;82603be06a8142aad1dfaa9e1279371f&lt;/code&gt;&lt;/p&gt;
&lt;h3&gt;
	Attempt 2 - reader B&lt;/h3&gt;
&lt;p&gt;
	Command line:&lt;/p&gt;
&lt;pre&gt;
&lt;code&gt;ddrescue -d -n -b 2048 /dev/sr0 windows_98_upgrade.iso windows_98_upgrade.log
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;
	Again this resulted in a 601.7 MB ISO image, again with no indication of read errors in the ddrescue log. MD5 checksum was (again) &lt;code&gt;82603be06a8142aad1dfaa9e1279371f&lt;/code&gt;.&lt;/p&gt;
&lt;p&gt;
	Then by chance I discovered some text files in the image file weren’t readable, so I did a third try, now again with reader A.&lt;/p&gt;
&lt;h3&gt;
	Attempt 3 - reader A&lt;/h3&gt;
&lt;p&gt;
	Command line:&lt;/p&gt;
&lt;pre&gt;
&lt;code&gt;ddrescue -d -n -b 2048 /dev/sr0 windows_98_upgrade_test.iso windows_98_upgrade_test.log
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;
	This resulted in a 660.9 MB ISO file. Again no errors in the ddrescue log; MD5 checksum is &lt;code&gt;24f0f746d0817121253c6b1242d4246e&lt;/code&gt;. After mounting the image, the text files that were problematic in the earlier images were normally readable.&lt;/p&gt;
&lt;h3&gt;
	Attempt 4 - reader B&lt;/h3&gt;
&lt;p&gt;
	Command line:&lt;/p&gt;
&lt;pre&gt;
&lt;code&gt;ddrescue -d -n -b 2048 /dev/sr0 windows_98_upgrade_refurbished_onemoretry.iso windows_98_upgrade_refurbished_onemoretry.log
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;
	Result was identical to result of attempt 3!&lt;/p&gt;
&lt;p&gt;
	So summarising, 2 runs of ddrescue (using 2 different USB readers) resulted in exactly the same error, whereas the remaining 2 runs (again using 2 different readers) completed fine. So what’s going on here!?&lt;/p&gt;
&lt;h2&gt;
	Md5sum directly on physical CD&lt;/h2&gt;
&lt;p&gt;
	As a first step I computed the MD5 checksum directly on the phyical disc, using:&lt;/p&gt;
&lt;pre&gt;
&lt;code&gt;md5sum /dev/sr0
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;
	I repeated this 4 times, using both readers A and B, plugging them into different USB slots. In each case the result was &lt;code&gt;24f0f746d0817121253c6b1242d4246e&lt;/code&gt;, which is identical to the hash I got for the ISO in attempts 3 and 4 (confirming these images are correct).&lt;/p&gt;
&lt;h2&gt;
	Comparison of ISO images in hex editor&lt;/h2&gt;
&lt;p&gt;
	I also did a comparison of the intact and faulty ISOs in a hex editor. This revealed that in the faulty images a block of about 59 MB of data is missing at the end of the file. I double checked this by copying the block of missing data to a separate file (missingblock.dat), after which I appended it to one of the faulty files using:&lt;/p&gt;
&lt;pre&gt;
&lt;code&gt;cat windows_98_upgrade_nl.iso missingblock.dat &amp;gt; isorepaired.iso
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;
	Then check:&lt;/p&gt;
&lt;pre&gt;
&lt;code&gt;md5sum isorepaired.iso
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;
	Result:&lt;/p&gt;
&lt;pre&gt;
&lt;code&gt;24f0f746d0817121253c6b1242d4246e
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;
	Which corresponds to the value of the intact image.&lt;/p&gt;
&lt;h2&gt;
	But why is this happening in the first place?!&lt;/h2&gt;
&lt;p&gt;
	The really important question is why this is happening in the first place, and if there’s any way to avoid it? The thread below on the ddrescue mailing list describes a somewhat similar (but not quite the same) issue:&lt;/p&gt;
&lt;p&gt;
	&lt;a href=&quot;https://lists.gnu.org/archive/html/bug-ddrescue/2014-02/msg00003.html&quot; rel=&quot;nofollow&quot;&gt;https://lists.gnu.org/archive/html/bug-ddrescue/2014-02/msg00003.html&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;
	Note the following quote from the &lt;a href=&quot;https://lists.gnu.org/archive/html/bug-ddrescue/2014-02/msg00009.html&quot; rel=&quot;nofollow&quot;&gt;response&lt;/a&gt; by ddrescue’s main author. He suggests that the problem &lt;em&gt;might&lt;/em&gt; an issue with a USB port, adding:&lt;/p&gt;
&lt;blockquote&gt;
	&lt;p&gt;
		Ddrescue can’t know if the data are really good or if the hardware is lying about it.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;
	If correct, this would apply to other imaging tools as well. Based on my results, I’m curious if other people may have run into similar issues. More importantly: how does one even detect errors like these? Of course it is always possible to run a checksum on the physical medium and then compare it to the ISO checksum, but this takes ages. A more quick and dirty approach would be to compare the size of each created image against the size of input medium. E.g. to get the size of a CD-ROM I can use something like this:&lt;/p&gt;
&lt;pre&gt;
&lt;code&gt;lsblk /dev/sr0 -n -b 
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;
	Result:&lt;/p&gt;
&lt;pre&gt;
&lt;code&gt;sr0   11:0    1 660850688  0 rom
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;
	(Third column is size of CD in bytes).&lt;/p&gt;
&lt;p&gt;
	To get the size of the ISO image:&lt;/p&gt;
&lt;pre&gt;
&lt;code&gt;du -b windows_98_upgrade_refurbished_onemoretry.iso
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;
	Result:&lt;/p&gt;
&lt;pre&gt;
&lt;code&gt;660850688   windows_98_upgrade_refurbished_onemoretry.iso
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;
	This does not guarantee the image is correct, but it will detect missing blocks of data.&lt;/p&gt;
&lt;p&gt;
	I also ran some cursory checks with &lt;em&gt;isovfy&lt;/em&gt; and &lt;em&gt;isoinfo&lt;/em&gt;, but the output of those tools turned out to be identical for both faulty and intact images, so they’re probably not very helpful for this sort of error.&lt;/p&gt;
&lt;p&gt;
	I’m curious how other people/memory institutions are dealing with this. Any thoughts / suggestions are welcome!&lt;/p&gt;
&lt;h2&gt;
	Addition&lt;/h2&gt;
&lt;p&gt;
	On Twitter Alexander Duryee rightly pointed out that a CD-Rom's Primary Volume Descriptor contains a field with the size of the disk (this is also where lsblk gets this value). So one would assume that ddrescue would check against this number. Apparently it doesn't do this, so I think I'll reprt this as a bug. (Note that such a check doesn't guarantee the copied data are identical to the source disc.)&lt;/p&gt;</description>
<guid isPermaLink="true">https://qanda.digipres.org/1076/incomplete-image-after-imaging-rom-prevent-and-detect-this</guid>
<pubDate>Thu, 03 Sep 2015 10:33:26 +0000</pubDate>
</item>
<item>
<title>How to detect Hybrid ISOs?</title>
<link>https://qanda.digipres.org/521/how-to-detect-hybrid-isos</link>
<description>&lt;p&gt;
	Asking for a friend:&lt;/p&gt;
&lt;p&gt;
	How does one detect a hybrid file system ISO, the ones that have HFS and Joliet for example?&lt;/p&gt;
&lt;p&gt;
	&lt;span style=&quot;font-family:courier new,courier,monospace;&quot;&gt;$ head -c 2048 my.iso | grep -c Apple_partition_map&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;
	Is there any tool that would give more meaningful metadata than &quot;ISO9660'?&lt;/p&gt;</description>
<guid isPermaLink="true">https://qanda.digipres.org/521/how-to-detect-hybrid-isos</guid>
<pubDate>Tue, 04 Nov 2014 15:37:39 +0000</pubDate>
</item>
<item>
<title>When to redact AIPs vs DIPs in digital archives workflows?</title>
<link>https://qanda.digipres.org/121/when-to-redact-aips-vs-dips-in-digital-archives-workflows</link>
<description>&lt;p&gt;
	Tools like bitcurator are integrating processes for &lt;a rel=&quot;nofollow&quot; href=&quot;http://wiki.bitcurator.net/index.php?title=BitCurator_and_Archival_Workflows&quot;&gt;redacting information from disk images&lt;/a&gt;. So, not just identifying information that you might want to use to make aprasial decisions about what to keep but also the ability to directly redact information from files. I know some folks are planning on using those tools that auto locate personal info and then redacting files or redacting information from inside files.&lt;/p&gt;
&lt;p&gt;
	Given that we have (at least in theory) Archival Information Packages and Disimenation Information packages we could be &lt;span class=&quot;il&quot;&gt;redacting&lt;/span&gt;, when should archivists be &lt;span class=&quot;il&quot;&gt;redacting&lt;/span&gt; the archival copy and when do we want to just be &lt;span class=&quot;il&quot;&gt;redacting&lt;/span&gt; the access materials? If one has information that is sensitive now but would be useful and non-sensitive in 50 years? Or, is the threat of mantaining sensitive information and multiple copies of materials too significant to warrent such an approach? Or, is this something that is going to depend on the particular issues in a collection. If the answer is &quot;it depends&quot; what is it that it depends on?&lt;/p&gt;</description>
<guid isPermaLink="true">https://qanda.digipres.org/121/when-to-redact-aips-vs-dips-in-digital-archives-workflows</guid>
<pubDate>Tue, 24 Jun 2014 14:22:05 +0000</pubDate>
</item>
</channel>
</rss>