<?xml version="1.0" encoding="UTF-8"?>
<rss xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:rdf="http://www.w3.org/1999/02/22-rdf-syntax-ns#" xmlns:taxo="http://purl.org/rss/1.0/modules/taxonomy/" version="2.0">
  <channel>
    <title>topic Re: How can i index scanned document imagefiles based on content in Alfresco Archive</title>
    <link>https://connect.hyland.com/t5/alfresco-archive/how-can-i-index-scanned-document-imagefiles-based-on-content/m-p/47476#M26940</link>
    <description>&lt;HTML&gt;&lt;HEAD&gt;&lt;/HEAD&gt;&lt;BODY&gt;&lt;SPAN&gt;Hai,&lt;/SPAN&gt;&lt;BR /&gt;&lt;SPAN&gt;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp; I have installed Alfresco 4.2 version in my system. I have added some documents in it and it will be listed. But how can i indexing all those documents. Pls anybody have an idea replay this. And also how to integrate OCR with this 4.2 version.&lt;/SPAN&gt;&lt;BR /&gt;&lt;BR /&gt;&lt;SPAN&gt;Thanks &lt;/SPAN&gt;&lt;BR /&gt;&lt;BR /&gt;&lt;SPAN&gt; Syed&lt;/SPAN&gt;&lt;BR /&gt;&lt;/BODY&gt;&lt;/HTML&gt;</description>
    <pubDate>Thu, 09 Jan 2014 07:22:32 GMT</pubDate>
    <dc:creator>syedmeeran</dc:creator>
    <dc:date>2014-01-09T07:22:32Z</dc:date>
    <item>
      <title>How can i index scanned document imagefiles based on content</title>
      <link>https://connect.hyland.com/t5/alfresco-archive/how-can-i-index-scanned-document-imagefiles-based-on-content/m-p/47470#M26934</link>
      <description>Hi all,&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp; Wht i understood abt indexing of pdf files when compared with normal text files is that it will use the text file generated from pdf file conversion for indexing purpose.But i have scanned image documents of type tiff files.&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp; My doubt is in this situation hw can i manage indexing</description>
      <pubDate>Tue, 19 Sep 2006 08:07:34 GMT</pubDate>
      <guid>https://connect.hyland.com/t5/alfresco-archive/how-can-i-index-scanned-document-imagefiles-based-on-content/m-p/47470#M26934</guid>
      <dc:creator>kishore</dc:creator>
      <dc:date>2006-09-19T08:07:34Z</dc:date>
    </item>
    <item>
      <title>Re: How can i index scanned document imagefiles based on content</title>
      <link>https://connect.hyland.com/t5/alfresco-archive/how-can-i-index-scanned-document-imagefiles-based-on-content/m-p/47471#M26935</link>
      <description>&lt;HTML&gt;&lt;HEAD&gt;&lt;/HEAD&gt;&lt;BODY&gt;&lt;SPAN&gt;Hi,&lt;/SPAN&gt;&lt;BR /&gt;&lt;BR /&gt;&lt;SPAN&gt;I too have a similar problem to that you mentioned. I have a buch of tiffs which will be stored in the repository, but I cant do a content search on these documents as there is no search index provided for this. I hope, right now we can do only search on the document name.&lt;/SPAN&gt;&lt;BR /&gt;&lt;BR /&gt;&lt;SPAN&gt;You are right that the documents will be indexed by converting into text and the indecies will have pointers to the documents , which will be used by a search query.&lt;/SPAN&gt;&lt;BR /&gt;&lt;BR /&gt;&lt;SPAN&gt;Probably we need to look in a persepctive to find a repository service that provides the indexing service and link to a document.&amp;nbsp; &lt;/SPAN&gt;&lt;BR /&gt;&lt;BR /&gt;&lt;SPAN&gt;I will update the thread if found a solution on this.&lt;/SPAN&gt;&lt;BR /&gt;&lt;BR /&gt;&lt;SPAN&gt;Any help from others is appreciated.&lt;/SPAN&gt;&lt;BR /&gt;&lt;SPAN&gt; &lt;/SPAN&gt;&lt;BR /&gt;&lt;SPAN&gt;Cheers,&lt;/SPAN&gt;&lt;BR /&gt;&lt;SPAN&gt;Pavan Kumar&lt;/SPAN&gt;&lt;/BODY&gt;&lt;/HTML&gt;</description>
      <pubDate>Tue, 19 Sep 2006 14:03:53 GMT</pubDate>
      <guid>https://connect.hyland.com/t5/alfresco-archive/how-can-i-index-scanned-document-imagefiles-based-on-content/m-p/47471#M26935</guid>
      <dc:creator>pavan_kumar</dc:creator>
      <dc:date>2006-09-19T14:03:53Z</dc:date>
    </item>
    <item>
      <title>Re: How can i index scanned document imagefiles based on content</title>
      <link>https://connect.hyland.com/t5/alfresco-archive/how-can-i-index-scanned-document-imagefiles-based-on-content/m-p/47472#M26936</link>
      <description>&lt;HTML&gt;&lt;HEAD&gt;&lt;/HEAD&gt;&lt;BODY&gt;&lt;SPAN&gt;Hi !&lt;/SPAN&gt;&lt;BR /&gt;&lt;BR /&gt;&lt;SPAN&gt;There are several OCR software which can be linked to Alfresco.&lt;/SPAN&gt;&lt;BR /&gt;&lt;SPAN&gt;For example : &lt;/SPAN&gt;&lt;A href="http://wiki.alfresco.com/wiki/Tiger_OCR_integration" rel="nofollow noopener noreferrer"&gt;http://wiki.alfresco.com/wiki/Tiger_OCR_integration&lt;/A&gt;&lt;BR /&gt;&lt;SPAN&gt;and other, but I can't found the web page.&lt;/SPAN&gt;&lt;BR /&gt;&lt;BR /&gt;&lt;SPAN&gt;Cheers,&lt;/SPAN&gt;&lt;BR /&gt;&lt;BR /&gt;&lt;SPAN&gt;Sam&lt;/SPAN&gt;&lt;/BODY&gt;&lt;/HTML&gt;</description>
      <pubDate>Tue, 19 Sep 2006 14:11:36 GMT</pubDate>
      <guid>https://connect.hyland.com/t5/alfresco-archive/how-can-i-index-scanned-document-imagefiles-based-on-content/m-p/47472#M26936</guid>
      <dc:creator>sam69</dc:creator>
      <dc:date>2006-09-19T14:11:36Z</dc:date>
    </item>
    <item>
      <title>Re: How can i index scanned document imagefiles based on content</title>
      <link>https://connect.hyland.com/t5/alfresco-archive/how-can-i-index-scanned-document-imagefiles-based-on-content/m-p/47473#M26937</link>
      <description>&lt;HTML&gt;&lt;HEAD&gt;&lt;/HEAD&gt;&lt;BODY&gt;&lt;SPAN&gt;Hi&lt;/SPAN&gt;&lt;BR /&gt;&lt;BR /&gt;&lt;SPAN&gt;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp; What i mean is already using other OCR software then hw can i link it?&lt;/SPAN&gt;&lt;BR /&gt;&lt;BR /&gt;&lt;SPAN&gt;Thanks&lt;/SPAN&gt;&lt;BR /&gt;&lt;SPAN&gt;kishore&lt;/SPAN&gt;&lt;/BODY&gt;&lt;/HTML&gt;</description>
      <pubDate>Thu, 21 Sep 2006 03:59:17 GMT</pubDate>
      <guid>https://connect.hyland.com/t5/alfresco-archive/how-can-i-index-scanned-document-imagefiles-based-on-content/m-p/47473#M26937</guid>
      <dc:creator>kishore</dc:creator>
      <dc:date>2006-09-21T03:59:17Z</dc:date>
    </item>
    <item>
      <title>Re: How can i index scanned document imagefiles based on content</title>
      <link>https://connect.hyland.com/t5/alfresco-archive/how-can-i-index-scanned-document-imagefiles-based-on-content/m-p/47474#M26938</link>
      <description>&lt;HTML&gt;&lt;HEAD&gt;&lt;/HEAD&gt;&lt;BODY&gt;&lt;SPAN&gt;I see at least two ways of achieving your goal:&lt;/SPAN&gt;&lt;BR /&gt;&lt;BR /&gt;&lt;SPAN&gt;1) build an Alfresco Package Archive, which is in fact a zip file containing your PDF files + 1 file that defines the meta-data. You will have to write a converter between your "custom" index file format and Alfresco XML. Try to export a space and look at the output to get an idea&lt;/SPAN&gt;&lt;BR /&gt;&lt;BR /&gt;&lt;SPAN&gt;2)&amp;nbsp; create a space and define a custom action, which parses index files (.txt in your case?), reads all the referenced PDF files (that you will have put in the same space before) and attach the meta-data "on-the-fly"&lt;/SPAN&gt;&lt;BR /&gt;&lt;BR /&gt;&lt;SPAN&gt;Solution 1 is more elegant and you have more control in case of errors. Solution 2 is easier, take a look at Alfresco SDK "Custom Action" project to start writing custom actions.&lt;/SPAN&gt;&lt;BR /&gt;&lt;BR /&gt;&lt;SPAN&gt;David&lt;/SPAN&gt;&lt;/BODY&gt;&lt;/HTML&gt;</description>
      <pubDate>Mon, 25 Sep 2006 16:33:13 GMT</pubDate>
      <guid>https://connect.hyland.com/t5/alfresco-archive/how-can-i-index-scanned-document-imagefiles-based-on-content/m-p/47474#M26938</guid>
      <dc:creator>dschmalz</dc:creator>
      <dc:date>2006-09-25T16:33:13Z</dc:date>
    </item>
    <item>
      <title>Re: How can i index scanned document imagefiles based on content</title>
      <link>https://connect.hyland.com/t5/alfresco-archive/how-can-i-index-scanned-document-imagefiles-based-on-content/m-p/47475#M26939</link>
      <description>&lt;HTML&gt;&lt;HEAD&gt;&lt;/HEAD&gt;&lt;BODY&gt;&lt;SPAN&gt;i just got approach 1 working for OCR for AnyDoc 3.2 on Alfresco 2.0. i use it to import medical claims images. here's how it works:&lt;/SPAN&gt;&lt;BR /&gt;&lt;BR /&gt;&lt;SPAN&gt;1. a script observes the AnyDoc output directory, when it finds a TXT output file it reads the file, counts the records, and compares the count to the number/names of images in the image output directory.&lt;/SPAN&gt;&lt;BR /&gt;&lt;SPAN&gt;2. if these match, i read in an XML template (based on the ACP XML schema for my custom claims content type) and replace it's values with the values from the AnyDoc output (once for each image).&lt;/SPAN&gt;&lt;BR /&gt;&lt;SPAN&gt;3. then I write out the XML, copy over the image files, zip the whole thing up into an ACP, and move the ACP into a CIFS directoy that's got an action to import an ACP to my claims space.&lt;/SPAN&gt;&lt;BR /&gt;&lt;BR /&gt;&lt;SPAN&gt;Right now the whole thing is just a proof-of-concept. I'd like to move the whole process into the Alfresco environment and then figure out a robust way to handle and report errors.&lt;/SPAN&gt;&lt;/BODY&gt;&lt;/HTML&gt;</description>
      <pubDate>Wed, 11 Apr 2007 13:50:05 GMT</pubDate>
      <guid>https://connect.hyland.com/t5/alfresco-archive/how-can-i-index-scanned-document-imagefiles-based-on-content/m-p/47475#M26939</guid>
      <dc:creator>johnwehall</dc:creator>
      <dc:date>2007-04-11T13:50:05Z</dc:date>
    </item>
    <item>
      <title>Re: How can i index scanned document imagefiles based on content</title>
      <link>https://connect.hyland.com/t5/alfresco-archive/how-can-i-index-scanned-document-imagefiles-based-on-content/m-p/47476#M26940</link>
      <description>&lt;HTML&gt;&lt;HEAD&gt;&lt;/HEAD&gt;&lt;BODY&gt;&lt;SPAN&gt;Hai,&lt;/SPAN&gt;&lt;BR /&gt;&lt;SPAN&gt;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp; I have installed Alfresco 4.2 version in my system. I have added some documents in it and it will be listed. But how can i indexing all those documents. Pls anybody have an idea replay this. And also how to integrate OCR with this 4.2 version.&lt;/SPAN&gt;&lt;BR /&gt;&lt;BR /&gt;&lt;SPAN&gt;Thanks &lt;/SPAN&gt;&lt;BR /&gt;&lt;BR /&gt;&lt;SPAN&gt; Syed&lt;/SPAN&gt;&lt;BR /&gt;&lt;/BODY&gt;&lt;/HTML&gt;</description>
      <pubDate>Thu, 09 Jan 2014 07:22:32 GMT</pubDate>
      <guid>https://connect.hyland.com/t5/alfresco-archive/how-can-i-index-scanned-document-imagefiles-based-on-content/m-p/47476#M26940</guid>
      <dc:creator>syedmeeran</dc:creator>
      <dc:date>2014-01-09T07:22:32Z</dc:date>
    </item>
  </channel>
</rss>

