<!DOCTYPE TEI.2 PUBLIC '-//C. M. Sperberg-McQueen//DTD
          TEI Lite 1.0 plus SWeb (XML)//EN'
          '../../../lib/swebxml.dtd' [
<!ATTLIST list type CDATA 'bullets' >
<!ATTLIST seg  rend CDATA 'incremental' >
<!ATTLIST xref href CDATA '' >

<!ATTLIST item id ID #IMPLIED >

<!ENTITY date.last.touched '3 August 2011'>

<!ENTITY eacute  "&#233;" ><!-- small e, acute accent -->

<!--* direct clone of 2010/12/idcc, with modest changes
    * to pitch it toward a CS audience not librarians
    *-->
<!ENTITY and    "&#x2227;" ><!--/wedge /land B: =logical and-->
<!ENTITY implies "&rArr;" ><!--/Rightarrow A: =implies-->
<!ENTITY mdash  "&#x2014;" ><!--=em dash-->
<!ENTITY or     "&#x2228;" ><!--/vee /lor B: =logical or-->
<!ENTITY rArr   "&#x21D2;" ><!--/Rightarrow A: =implies-->

<!ENTITY larr   "&#x2190;" ><!--/leftarrow /gets A: =leftward arrow-->
<!ENTITY rarr   "&#x2192;" ><!--/rightarrow /to A: =rightward arrow-->
<!ENTITY uarr   "&#x2191;" ><!--/uparrow A: =upward arrow-->
<!ENTITY darr   "&#x2193;" ><!--/downarrow A: =downward arrow-->

]>
<?xml-stylesheet type="text/xsl" href="local.xsl"?> 
<TEI.2>
  <teiHeader>
    <fileDesc>
      <titleStmt>
	<title>The Model Editions Partnership</title>
	<title>Fifteen yeare on</title>
      </titleStmt>
      <publicationStmt>
	<p>Unpublished and confidential.</p>
      </publicationStmt>
      <sourceDesc>
	<p>Created in electronic form.</p>
      </sourceDesc>
    </fileDesc>
    <profileDesc>
      <textClass>
	<classCode>slides</classCode>
      </textClass>
    </profileDesc>
  </teiHeader>
  <text>
    <front>
      <titlePage>
	<docTitle>
	  <titlePart>The Model Editions Partnership</titlePart>
	  <titlePart>Fifteen years on</titlePart>
	</docTitle>

	<titlePart>Balisage 2011</titlePart>
	<docDate>Montr&eacute;al, &date.last.touched;</docDate>
	<titlePart><xref href="http://www.blackmesatech.com/2011/08/mep/"
			 >http://www.blackmesatech.com/2011/08/mep/</xref></titlePart>
      </titlePage>

</front>

<body>

<!--*
      <div>
	<head>Preliminaries</head>
	<list>
	  <item>a proposition</item>
	  <item>a view from outside</item>
	</list>
      </div>
*-->

      <div>
	<head>Overview</head>
	<list>
	  <item>The data longevity promise (and a challenge)</item>
	  <item>The Model Editions Partnership</item>
	  <item>A report card</item>
	  <item>Lessons for the future</item>
	</list>
      </div>

      <div>
	<head>The data longevity promise</head>
	<p>Coombs, Renear, DeRose (1987):
	  <q type="block"><p>
	      ... an author who is not using descriptive markup
	      may have to modify the markup of document files
	      <!--* for any of several reasons: ... *-->
	      ...
	      whenever the author or the
	      installation changes the computing environment. 
	      <!--*
	      When FRESS (File Retrieval and Editing System) users 
	      at Brown University learned FRESS would no longer be 
	      supported, authors either spent hours converting their 
	      files to the new format (Waterloo SCRIPT) or accepted 
	      the possibility of <soCalled>losing</soCalled> their 
	      *-->
	      When [systems are replaced], authors either [spend] 
	      hours converting their 
	      files to the new format ... or [accept] 
	      the possibility of <soCalled>losing</soCalled> their 
	      data. <!--* Even updates in the current text formatter 
	      often require modifications in files. Changing to a 
	      new printer may require modifications. *-->
	      ... In fact, almost 
	      any change in the computing environment poses a threat 
	      to one’s files if they contain procedural or 
	      presentational markup.
	  </p></q>
	</p>
	<p>But in contrast
	<q type="block"><p>As we point out [elsewhere], 
	descriptive markup eliminates maintenance concerns. </p></q>
	</p>
      </div>

      <div>
	<head>Twenty years later, ...</head>
	<p>One disadvantage of clear statements:
	they are falsifiable.</p>
      </div>

      <div>
	<head>The Model Editions Partnership</head>
	<p>Multi-stage, multi-year project.</p>
	<p><list>
	  <item><label>funded by</label> 
	  U.S. National Historical Publications and Records Commission 
	  (NHPRC)</item>
	  <item><label>based at</label> 
	  University of South Carolina</item>
	  <item><label>led by</label> 
	  David R. Chesnutt</item>
	</list></p>
      </div>
      <div>
      <head>MEP I</head>
	<p>Original project (1995-1999):  
	<list>
	  <item><xref href="http://wyatt.elasticbeanstalk.com/mep/misc/prospectus.html"
		      ><title>A Prospectus for Electronic Historical Editions</title></xref> (1996)</item>
	  <item>Adaptation of TEI P3 for historical documentary editions
	  (<xref href="http://wyatt.elasticbeanstalk.com/mep/misc/meptsd2.2.html"
	  >TEI Tag Set Description</xref>)</item>
	  <item>Web site (<xref href="http://mep.cla.sc.edu/">http://mep.cla.sc.edu/</xref>)
	  with samples (150-200 pages each) from seven partner projects:
	  <list>
	    <item>Documentary History of the First Federal Congress</item>
	    <item>Documentary History of the Ratification of the Constitution and the Bill of Rights </item>
	    <item>Papers of General Nathanael Greene </item>
	    <item>Papers of Henry Laurens </item>
	    <item>Lincoln Legal Papers</item>
	    <item>Margaret Sanger Papers </item>
	    <item>Papers of Elizabeth Cady Stanton and Susan B. Anthony </item>
	  </list>
	  </item>
	</list>
	</p>
      </div>
      <div>
      <head>MEP II etc.</head>
	<p>Later work:
	<list>
	  <item>tweaks to the DTD</item>
	  <item>XML version of the DTD</item>
	  <item>more samples
	  <list>
	    <item>The Frederick Douglass Papers  </item>
	    <item>The Papers of Dwight David Eisenhower  </item>
	    <item>The Marcus Garvey and UNIA Papers  </item>
	    <item>The Papers of Joseph Henry  </item>
	    <item>The Papers of George Catlett Marshall  </item>
	    <item>Eleanor Roosevelt Papers </item>
	  </list></item>
	  <item>stand-alone deliverable on CD-ROM</item>
	</list>
	</p>
      </div>
      <div>
      <head>Time spares no software</head>
	<p>2009:  the server at USC dies.</p>
	<p>2010-2011:  attempts to re-host the material.
	Currently in the cloud 
	(<xref href="http://mep.blackmesatech.com/mep/"
	>http://mep.blackmesatech.com/mep/</xref>).
	</p>
      </div>

      <div>
	<head>So ... how'd we do?</head>
	<p>Did descriptive markup buffer us from changes
	in the environment?</p>
	<p>Did it eliminate maintenance concerns?</p>
	<p>What worked?  What didn't work?</p>
	<p>If we had it to do again, ...?</p>
      </div>

      <div>
	<head>Keepers</head>
	<p>Things that worked / if you were doing it again today,
	what would you do in the same way?</p>
      </div>
      <div>
	<head>Use descriptive markup</head>
	<p>Yes, it works.</p>
	<p>No, it's not flawless.*</p>
      </div>
      <div>
	<head>Use TEI</head>
	<p>Because the analysis works fairly well.</p>
	<p><emph>Not</emph> for reuse of stylesheets etc.</p>
      </div>
      <div>
	<head>Modify TEI</head>
	<p>Because modifications allow a better fit.</p>
      </div>
      <div>
	<head>Multiple vocabulary flavors</head>
	<p>MEP's design distinguishes
	<list>
	  <item><soCalled>gold-standard</soCalled>
	  archival vocabulary (for long-term storage,
	  convenient delivery)</item>
	  <item>data-capture vocabulary (lighter-weight
	  structure and metadata)</item>
	</list>
	</p>
      </div>
      <div>
	<head>Deliver early and often</head>
	<p>MEP had some skeptical partners (and funders).</p>
	<p>So there was pressure to have <emph>something
	nice to show people</emph>.</p>
      </div>
      <div>
	<head>Big books and little books</head>
	<p>MEP's practice* distinguishes
	<list>
	  <item><soCalled>big-book</soCalled> encoding</item>
	  <item><soCalled>little-book</soCalled> encoding</item>
	</list>
	* The practice, but not the theory ... (why not?)
	</p>
      </div>
      <div>
	<head>Assign identifiers</head>
	<p>MEP's practice assigns identifiers to each
	singly deliverable item
	(document, biographical note, major section
	of introductory essay).
	</p>
	<p>We may have done this for the wrong reasons.</p>
	<p rend="incremental">But it's the right thing to do.</p>
      </div>

      <div>
	<head>Lessons for the future</head>
	<p>What would you do differently if you were doing it today?</p>
      </div>

      <div>
	<head>Use current standards and software</head>
	<list>
	  <item>XML, not SGML</item>
	  <item>XSLT</item>
	  <item>XQuery</item>
	  <item>XProc</item>
	  <item>XForms</item>
	</list>
	<p>This is <q>different</q> in detail but not substance.</p>
      </div>

      <div>
	<head>Serve the XML directly</head>
	<p>More feasible today (but Panorama).</p>
      </div>

      <div>
	<head>Don't fake small caps</head>
	<p>In 1999, browsers didn't do small caps at all (let alone well).</p>
	<p>But this is not the answer:
	<eg><![CDATA[
	<p TEIform="p">In committee of the whole 
          on the State of the Union, 
          <person reg="Trumbull, Jonathan" TEIform="name"
          >Mr. T<sCap TEIform="hi">RUMBULL</sCap></person> 
          in the chair.</p>
  	]]></eg>
	</p>
      </div>

      <div>
	<head>Make the archival form</head>
	<p>MEP's design distinguishes
	<list>
	  <item><soCalled>gold-standard</soCalled>
	  archival vocabulary (for long-term storage,
	  convenient delivery)</item>
	  <item>data-capture vocabulary (lighter-weight
	  structure and metadata)</item>
	</list>
	</p>
	<p>But we never actually got around to making
	the archival form.</p>
	<p>Delivery system was built around data-capture 
	format.</p>
      </div>
      <div>
	<head>Hire a metadata Nazi</head>
	<p>The documents do have change logs, but
	they are almost entirely useless.
	<eg><![CDATA[
<doc id="fc10720">
  <mepHeader>
    <prepDate>cm 23 Jan. 1997; QA and validation by LG/ 24 Jan. 9</prepDate>
    <prepDate>2/24/97; QA CM- 2/16/97</prepDate>
    <prepDate>9 April 97, Level 3; QA 4/11/97</prepDate>
    <prepDate>rg 97-05-13; QA 97-05-20, cm</prepDate>
    <prepDate>97-08-28 rg</prepDate>
    <prepDate>97-09-26 ml </prepDate>
    <prepDate> 98-02-05 ah </prepDate>
    <prepDate>98-02-10 rg </prepDate>
    <prepDate>98-05-27 mm</prepDate>
    <prepDate>98-07-06 ml </prepDate>
    <prepDate>98-07-15 rg</prepDate>
    <prepDate>99-04-05 ah</prepDate>
    <idno TEIform="idno">FC10720</idno>
    <docTitle TEIform="docTitle">
      <titlePart type="main" TEIform="titlePart">
      <title TEIform="title">Gazette of the United States</title>, 20 May 1789</titlePart>
    </docTitle>
    ...
      ]]>
      </eg></p>
      </div>
      <div>
	<head>Have a toolmaker around</head>
	<p>Even today, projects will have
	special needs hard to meet with
	off-the-shelf packages.</p>
      </div>
      <div>
	<head>Document processing and styles better</head>
	<p>The more effort you put into styling,
	the less you want to lose it.</p>
	<p>If you can't use standard languages, at least
	document what you did in prose.</p>
      </div>
      <div>      
	<head>Multiple-use</head>
	<p>Work harder to provide delivery using two
	different systems.</p>
	<p>Also revise your first version.  (Don't stand
	still.)</p>
      </div>

      <div>      
	<head>Memento mori.</head>

	<list>
	  <item>Routine backups (even if <q>nothing</q> has changed)</item>
	  <item>Document behaviors and appearance (?)</item>
	  <item>Have a successor lined up.</item>
	</list>
	<p>(All easier said than done.)</p>
      </div>


<!--*
	  <item>
   To the extent that you put a lot of intellectual effort into stylesheets
     and design of interaction, you need to do so in a reusable form,
     either by using standard stylesheet languages or by documenting
     carefully what you are doing.  Otherwise you end up with 
     partners unhappy because the look and feel of the new site
     are not the same, and you can't replicate the old look and feel
     because it's not documented and the software won't run.
     (Alternatively, train you partners away from the idea that there
     is or should be only one look and feel.  This is a good idea in
     principle, but hard to put into practice.)
	  </item>
	  <item>
	  
   Memento mori.  It's easier to re-locate a body of material if all
     its affairs are in order:  things are documented, there are current
     backups, and so on.  To the extent that that was the case with 
     MEP, making the new site was easy; to the extent that it wasn't,
     it has been hard.  [need concrete examples]
	  </item>
	  <item>
	  
     Possible horror stories about five-year projects whose funding
     was lost suddenly and whose staff had only a few days, or a few
     hours, warning before the building was padlocked.  In one case,
     the project's data survive in part today only because a couple
     people stole the hard disks out of their machines.  That was a 
     misdeed by the funding agency - - but it also shows that the 
     project did not have a succession plan in place.
</item>
	</list>
*-->

<!--*
      <div>
	<head>Lessons for the future</head>
	<p>What would you do differently if you were doing it today?</p>
	<list>
	  <head>We'd do it in XML, of course.</head>
	  <head>Get to archival form faster and earlier.</head>
	  <head>Hire a metadata Nazi.</head>
	  <head>Get better documentation of processing and processing
	  expectations.</head>
	  <head>Multiple-use.</head>
	  <head>Serve the XML.</head>
	  <head>Memento mori.</head>
	  <item>
   it's not enough to define a gold-standard archival form and
     plan to migrate data into it from the data-capture format;
     you need to migrate the data in actuality, not just in your
     plans.  Otherwise, you end up having to build the delivery
     system around the data capture format which is less
     regular and makes everything harder.  (Concrete example:
     copyright information.)
	  </item>
	  <item>
	  
   Early visible results are good; early visible results with plans to
     learn from experience and revise them periodically would, I
     think, be better.  Once you have your first version, don't stand
     still.
	  </item>
	  <item>
	  
   To the extent that you put a lot of intellectual effort into stylesheets
     and design of interaction, you need to do so in a reusable form,
     either by using standard stylesheet languages or by documenting
     carefully what you are doing.  Otherwise you end up with 
     partners unhappy because the look and feel of the new site
     are not the same, and you can't replicate the old look and feel
     because it's not documented and the software won't run.
     (Alternatively, train you partners away from the idea that there
     is or should be only one look and feel.  This is a good idea in
     principle, but hard to put into practice.)
	  </item>
	  <item>
	  
   Memento mori.  It's easier to re-locate a body of material if all
     its affairs are in order:  things are documented, there are current
     backups, and so on.  To the extent that that was the case with 
     MEP, making the new site was easy; to the extent that it wasn't,
     it has been hard.  [need concrete examples]
	  </item>
	  <item>
	  
     Possible horror stories about five-year projects whose funding
     was lost suddenly and whose staff had only a few days, or a few
     hours, warning before the building was padlocked.  In one case,
     the project's data survive in part today only because a couple
     people stole the hard disks out of their machines.  That was a 
     misdeed by the funding agency - - but it also shows that the 
     project did not have a succession plan in place.
</item>
	</list>
      </div>

*-->


      <div>
	<head>Conclusions</head>
	<p>What does this mean?</p>
      </div>

</body>


</text>
</TEI.2>