<?xml version="1.0" encoding="UTF-8"?>
<rss xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:rdf="http://www.w3.org/1999/02/22-rdf-syntax-ns#" xmlns:taxo="http://purl.org/rss/1.0/modules/taxonomy/" version="2.0">
  <channel>
    <title>topic PDF Indexing in Alfresco Archive</title>
    <link>https://connect.hyland.com/t5/alfresco-archive/pdf-indexing/m-p/110997#M77999</link>
    <description>&lt;HTML&gt;&lt;HEAD&gt;&lt;/HEAD&gt;&lt;BODY&gt;&lt;SPAN&gt;I am new to Alfresco… downloaded and installed just a few hours ago. It's up and running and I can login and upload files and create users.&lt;/SPAN&gt;&lt;BR /&gt;&lt;BR /&gt;&lt;SPAN&gt;The problem I am facing is that the PDF files I uploaded can only be searched by filename/title, while other files e.g. Word .doc files can be searched using words in the file.&lt;/SPAN&gt;&lt;BR /&gt;&lt;BR /&gt;&lt;SPAN&gt;So, if I have Raptors.doc and Raptors.pdf, and both contain the word "Toronto" in them, a search for "Raptors" retrieves both documents, but a search for "Toronto" retrieves just the .doc file, not the .pdf file.&lt;/SPAN&gt;&lt;BR /&gt;&lt;BR /&gt;&lt;SPAN&gt;As a side note, both the .pdf and the .doc files were created from the same source in Google Docs &amp;amp; Spreadsheet.&lt;/SPAN&gt;&lt;BR /&gt;&lt;BR /&gt;&lt;SPAN&gt;Am I missing something?&lt;/SPAN&gt;&lt;BR /&gt;&lt;BR /&gt;&lt;SPAN&gt;Thanks in advance.&lt;/SPAN&gt;&lt;BR /&gt;&lt;BR /&gt;&lt;SPAN&gt;____________________________________&lt;/SPAN&gt;&lt;BR /&gt;&lt;BR /&gt;&lt;SPAN&gt;added after initial post&lt;/SPAN&gt;&lt;BR /&gt;&lt;SPAN&gt;____________________________________&lt;/SPAN&gt;&lt;BR /&gt;&lt;BR /&gt;&lt;SPAN&gt;I noticed that PDFs under 500kB have no issues with indexing. However, for *large* PDFs (say above 3 MB, which isn't really large, I have PDFs well over 70MB), I get the following error message:&lt;/SPAN&gt;&lt;BR /&gt;&lt;BR /&gt;&lt;SPAN&gt;Metadata extraction failed: reader: ContentAccessor[ contentURL=store://C:\Alfresco\tomcat\temp\Alfresco\alfresco39438.upload, mimetype=application/pdf, size=12976429, encoding=UTF-8] extracter: org.alfresco.repo.content.metadata.PdfBoxMetadataExtracter@17349d7&lt;/SPAN&gt;&lt;/BODY&gt;&lt;/HTML&gt;</description>
    <pubDate>Mon, 30 Apr 2007 00:20:10 GMT</pubDate>
    <dc:creator>smca</dc:creator>
    <dc:date>2007-04-30T00:20:10Z</dc:date>
    <item>
      <title>PDF Indexing</title>
      <link>https://connect.hyland.com/t5/alfresco-archive/pdf-indexing/m-p/110997#M77999</link>
      <description>I am new to Alfresco… downloaded and installed just a few hours ago. It's up and running and I can login and upload files and create users.The problem I am facing is that the PDF files I uploaded can only be searched by filename/title, while other files e.g. Word .doc files can be searched using wor</description>
      <pubDate>Mon, 30 Apr 2007 00:20:10 GMT</pubDate>
      <guid>https://connect.hyland.com/t5/alfresco-archive/pdf-indexing/m-p/110997#M77999</guid>
      <dc:creator>smca</dc:creator>
      <dc:date>2007-04-30T00:20:10Z</dc:date>
    </item>
    <item>
      <title>Re: PDF Indexing</title>
      <link>https://connect.hyland.com/t5/alfresco-archive/pdf-indexing/m-p/110998#M78000</link>
      <description>&lt;HTML&gt;&lt;HEAD&gt;&lt;/HEAD&gt;&lt;BODY&gt;&lt;BLOCKQUOTE class="jive-quote"&gt;I noticed that PDFs under 500kB have no issues with indexing. However, for *large* PDFs (say above 3 MB, which isn't really large, I have PDFs well over 70MB), I get the following error message:&lt;BR /&gt;&lt;BR /&gt;Metadata extraction failed: reader: ContentAccessor[ contentURL=store://C:\Alfresco\tomcat\temp\Alfresco\alfresco39438.upload, mimetype=application/pdf, size=12976429, encoding=UTF-8] extracter: org.alfresco.repo.content.metadata.PdfBoxMetadataExtracter@17349d7&lt;/BLOCKQUOTE&gt;&lt;BR /&gt;&lt;SPAN&gt;Internally we use the open source PDFBox library to perform the to text conversion of PDF documents. It is possible the library is having a problem with certain documents - are there any other errors (stack trace?) in the log? Can you enter the following search in the search box in the web-client:&lt;/SPAN&gt;&lt;BR /&gt;&lt;BR /&gt;&lt;SPAN&gt;nift&lt;/SPAN&gt;&lt;BR /&gt;&lt;BR /&gt;&lt;SPAN&gt;it should return those documets that have failed to index due to the transformation engine failing.&lt;/SPAN&gt;&lt;BR /&gt;&lt;BR /&gt;&lt;SPAN&gt;If a transformation takes too long, it is shunted into a background thread, but it should still complete unless there is an error.&lt;/SPAN&gt;&lt;BR /&gt;&lt;BR /&gt;&lt;SPAN&gt;Thanks,&lt;/SPAN&gt;&lt;BR /&gt;&lt;BR /&gt;&lt;SPAN&gt;Kevin&lt;/SPAN&gt;&lt;/BODY&gt;&lt;/HTML&gt;</description>
      <pubDate>Mon, 30 Apr 2007 09:30:42 GMT</pubDate>
      <guid>https://connect.hyland.com/t5/alfresco-archive/pdf-indexing/m-p/110998#M78000</guid>
      <dc:creator>kevinr</dc:creator>
      <dc:date>2007-04-30T09:30:42Z</dc:date>
    </item>
    <item>
      <title>Re: PDF Indexing</title>
      <link>https://connect.hyland.com/t5/alfresco-archive/pdf-indexing/m-p/110999#M78001</link>
      <description>&lt;HTML&gt;&lt;HEAD&gt;&lt;/HEAD&gt;&lt;BODY&gt;&lt;SPAN&gt;I searched for nift - no results.&lt;/SPAN&gt;&lt;BR /&gt;&lt;BR /&gt;&lt;SPAN&gt;I searched for nitf - all the documents that failed indexing showed up!&lt;/SPAN&gt;&lt;BR /&gt;&lt;BR /&gt;&lt;SPAN&gt;I noticed that my setup can cope with large (i.e. &amp;gt;10MB) Word .doc and OpenOffice .sxw documents. It's having troubles only with .pdf files.&lt;/SPAN&gt;&lt;BR /&gt;&lt;BR /&gt;&lt;SPAN&gt;By the way, I am pretty new to Tomcat and Alfresco. How do I get a stack trace and where is the log file usually located?&lt;/SPAN&gt;&lt;/BODY&gt;&lt;/HTML&gt;</description>
      <pubDate>Mon, 30 Apr 2007 15:06:27 GMT</pubDate>
      <guid>https://connect.hyland.com/t5/alfresco-archive/pdf-indexing/m-p/110999#M78001</guid>
      <dc:creator>smca</dc:creator>
      <dc:date>2007-04-30T15:06:27Z</dc:date>
    </item>
    <item>
      <title>Re: PDF Indexing</title>
      <link>https://connect.hyland.com/t5/alfresco-archive/pdf-indexing/m-p/111000#M78002</link>
      <description>&lt;HTML&gt;&lt;HEAD&gt;&lt;/HEAD&gt;&lt;BODY&gt;&lt;SPAN&gt;Is there any way around the issue of indexing failures (at least in hardware if not in software)? Like adding more RAM (I have 2GB) or a more powerful processor (I am using and Athlon X2 6000+). As far as the specs of my machine is concerned, it is nothing to sneeze at. It's quite powerful.&lt;/SPAN&gt;&lt;BR /&gt;&lt;BR /&gt;&lt;SPAN&gt;At this time, I am still testing various document management systems (alfresco, dspace, plone, etc.) to keep track of all documents created by everybody in my family (currently over 5GB of .doc, .pdf, .xls, .ppt, .rtf … some are over 9 years old…). If hardware is the limiting factor on my "test" system, I would consider a more powerful system for the final deployment, like a quad-core CPU. How much more powerful can a home PC get beyond what I have in terms of CPU and RAM unless I move to 64-bit, which I may if push comes to shove.&lt;/SPAN&gt;&lt;/BODY&gt;&lt;/HTML&gt;</description>
      <pubDate>Tue, 08 May 2007 17:03:49 GMT</pubDate>
      <guid>https://connect.hyland.com/t5/alfresco-archive/pdf-indexing/m-p/111000#M78002</guid>
      <dc:creator>smca</dc:creator>
      <dc:date>2007-05-08T17:03:49Z</dc:date>
    </item>
    <item>
      <title>Re: PDF Indexing</title>
      <link>https://connect.hyland.com/t5/alfresco-archive/pdf-indexing/m-p/111001#M78003</link>
      <description>&lt;HTML&gt;&lt;HEAD&gt;&lt;/HEAD&gt;&lt;BODY&gt;&lt;SPAN&gt;The machine specs sound more than enough. The problem is the PDFBox library failing on certain PDF fails - it's nothing to do with memory/cpu given your machine specs. It may be worth a try trying to update the PDFBox library used in alfresco if there is a newer version available, otherwise we would need to submit bugs to the author. Failing that if you can find another PDF-&amp;gt;text java library then it could be configured in as the transformer instead of PDFBox.&lt;/SPAN&gt;&lt;BR /&gt;&lt;BR /&gt;&lt;SPAN&gt;Thanks,&lt;/SPAN&gt;&lt;BR /&gt;&lt;BR /&gt;&lt;SPAN&gt;Kevin&lt;/SPAN&gt;&lt;/BODY&gt;&lt;/HTML&gt;</description>
      <pubDate>Tue, 08 May 2007 17:27:04 GMT</pubDate>
      <guid>https://connect.hyland.com/t5/alfresco-archive/pdf-indexing/m-p/111001#M78003</guid>
      <dc:creator>kevinr</dc:creator>
      <dc:date>2007-05-08T17:27:04Z</dc:date>
    </item>
    <item>
      <title>Re: PDF Indexing</title>
      <link>https://connect.hyland.com/t5/alfresco-archive/pdf-indexing/m-p/111002#M78004</link>
      <description>&lt;HTML&gt;&lt;HEAD&gt;&lt;/HEAD&gt;&lt;BODY&gt;&lt;SPAN&gt;Environment: Alfresco 2.0, Tomcat5, Linux&lt;/SPAN&gt;&lt;BR /&gt;&lt;BR /&gt;&lt;SPAN&gt;When uploading a 3.3MB PDF the following error occures:&lt;/SPAN&gt;&lt;BR /&gt;&lt;BR /&gt;&lt;PRE class="language-none line-numbers"&gt;&lt;CODE&gt;14:51:10,569 WARN&amp;nbsp; [org.alfresco.web.bean.repository.Repository] Metadata extraction failed: &lt;BR /&gt;&amp;nbsp;&amp;nbsp; reader: ContentAccessor[ contentUrl=store:///srv/www/tomcat5/base/temp/Alfresco/alfresco40141.upload, mimetype=application/pdf, size=0, encoding=UTF-8]&lt;BR /&gt;&amp;nbsp;&amp;nbsp; extracter: org.alfresco.repo.content.metadata.PdfBoxMetadataExtracter@5cd46b&lt;BR /&gt;&lt;SPAN class="line-numbers-rows"&gt;&lt;SPAN&gt;‍&lt;/SPAN&gt;&lt;SPAN&gt;‍&lt;/SPAN&gt;&lt;SPAN&gt;‍&lt;/SPAN&gt;&lt;SPAN&gt;‍&lt;/SPAN&gt;&lt;/SPAN&gt;&lt;/CODE&gt;&lt;/PRE&gt;&lt;/BODY&gt;&lt;/HTML&gt;</description>
      <pubDate>Thu, 31 May 2007 12:59:30 GMT</pubDate>
      <guid>https://connect.hyland.com/t5/alfresco-archive/pdf-indexing/m-p/111002#M78004</guid>
      <dc:creator>manfred99</dc:creator>
      <dc:date>2007-05-31T12:59:30Z</dc:date>
    </item>
    <item>
      <title>Re: PDF Indexing</title>
      <link>https://connect.hyland.com/t5/alfresco-archive/pdf-indexing/m-p/111003#M78005</link>
      <description>&lt;HTML&gt;&lt;HEAD&gt;&lt;/HEAD&gt;&lt;BODY&gt;&lt;SPAN&gt;I've got a slightly different issue, in that I have .odt and .pdf files with exactly the same content (the .pdf is created from the .odt in OpenOffice).&amp;nbsp; When I run a search only the .pdf files are returned.&lt;/SPAN&gt;&lt;BR /&gt;&lt;BR /&gt;&lt;SPAN&gt;I am running v2.0.0 Community on SUSE 10.0, and have tried the search as several users including "guest" and "admin" and get the same result.&amp;nbsp; The documents are in the Guest Space, with consumer access to all.&lt;/SPAN&gt;&lt;/BODY&gt;&lt;/HTML&gt;</description>
      <pubDate>Thu, 31 May 2007 13:55:06 GMT</pubDate>
      <guid>https://connect.hyland.com/t5/alfresco-archive/pdf-indexing/m-p/111003#M78005</guid>
      <dc:creator>forcev</dc:creator>
      <dc:date>2007-05-31T13:55:06Z</dc:date>
    </item>
  </channel>
</rss>

