<?xml version="1.0" encoding="UTF-8"?>
<rss xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:rdf="http://www.w3.org/1999/02/22-rdf-syntax-ns#" xmlns:taxo="http://purl.org/rss/1.0/modules/taxonomy/" version="2.0">
  <channel>
    <title>topic Re: Indexing of PDF custom fields? in Alfresco Archive</title>
    <link>https://connect.hyland.com/t5/alfresco-archive/indexing-of-pdf-custom-fields/m-p/40009#M21351</link>
    <description>&lt;HTML&gt;&lt;HEAD&gt;&lt;/HEAD&gt;&lt;BODY&gt;&lt;SPAN&gt;Pascal,&lt;/SPAN&gt;&lt;BR /&gt;&lt;SPAN&gt;The same problem is in our system.&amp;nbsp; I've found it's much easier to deal with metadata in the file and have a way of extracting it afterwards, rather than always bringing the metadata with you in a separate database.&amp;nbsp; At least in Alfresco, Alfresco can pull out the necessary Title/description/author fields to display in your content view based on content rules.&lt;/SPAN&gt;&lt;BR /&gt;&lt;BR /&gt;&lt;SPAN&gt;PDF metadata extraction in Alfresco doesn't seem to work properly.&amp;nbsp; Sometimes nothing gets extracted, sometimes gibberish gets extracted. So I would be interested in a solution to this as you would.&lt;/SPAN&gt;&lt;/BODY&gt;&lt;/HTML&gt;</description>
    <pubDate>Mon, 23 Apr 2007 12:44:41 GMT</pubDate>
    <dc:creator>aatamer</dc:creator>
    <dc:date>2007-04-23T12:44:41Z</dc:date>
    <item>
      <title>Indexing of PDF custom fields?</title>
      <link>https://connect.hyland.com/t5/alfresco-archive/indexing-of-pdf-custom-fields/m-p/40008#M21350</link>
      <description>I would like to store scanned images in Alfresco, including some metadata (either recognized by OCR or manual indexing). We usually do this by creating a PDF, containing the metadata in custom fields (e.g. "Customer" =&amp;gt; "ACME Inc.").Problem: Alfresco doesn't seem to index the PDF custom fields… I</description>
      <pubDate>Thu, 13 Jul 2006 10:21:51 GMT</pubDate>
      <guid>https://connect.hyland.com/t5/alfresco-archive/indexing-of-pdf-custom-fields/m-p/40008#M21350</guid>
      <dc:creator>pascalsartorett</dc:creator>
      <dc:date>2006-07-13T10:21:51Z</dc:date>
    </item>
    <item>
      <title>Re: Indexing of PDF custom fields?</title>
      <link>https://connect.hyland.com/t5/alfresco-archive/indexing-of-pdf-custom-fields/m-p/40009#M21351</link>
      <description>&lt;HTML&gt;&lt;HEAD&gt;&lt;/HEAD&gt;&lt;BODY&gt;&lt;SPAN&gt;Pascal,&lt;/SPAN&gt;&lt;BR /&gt;&lt;SPAN&gt;The same problem is in our system.&amp;nbsp; I've found it's much easier to deal with metadata in the file and have a way of extracting it afterwards, rather than always bringing the metadata with you in a separate database.&amp;nbsp; At least in Alfresco, Alfresco can pull out the necessary Title/description/author fields to display in your content view based on content rules.&lt;/SPAN&gt;&lt;BR /&gt;&lt;BR /&gt;&lt;SPAN&gt;PDF metadata extraction in Alfresco doesn't seem to work properly.&amp;nbsp; Sometimes nothing gets extracted, sometimes gibberish gets extracted. So I would be interested in a solution to this as you would.&lt;/SPAN&gt;&lt;/BODY&gt;&lt;/HTML&gt;</description>
      <pubDate>Mon, 23 Apr 2007 12:44:41 GMT</pubDate>
      <guid>https://connect.hyland.com/t5/alfresco-archive/indexing-of-pdf-custom-fields/m-p/40009#M21351</guid>
      <dc:creator>aatamer</dc:creator>
      <dc:date>2007-04-23T12:44:41Z</dc:date>
    </item>
    <item>
      <title>Re: Indexing of PDF custom fields?</title>
      <link>https://connect.hyland.com/t5/alfresco-archive/indexing-of-pdf-custom-fields/m-p/40010#M21352</link>
      <description>&lt;HTML&gt;&lt;HEAD&gt;&lt;/HEAD&gt;&lt;BODY&gt;&lt;SPAN&gt;Hi Friends,&lt;/SPAN&gt;&lt;BR /&gt;&lt;BR /&gt;&lt;SPAN&gt;&amp;nbsp;&amp;nbsp; I'm also facing the same problem. Actually I'm able to read data from PDF (PDF forms) files using PDFBox but I want this information to be extracted as metadata just like the way Name,/Author/Desc of PDF document and it should be displayed on screen.&lt;/SPAN&gt;&lt;BR /&gt;&lt;BR /&gt;&lt;SPAN&gt;I have my own custom aspect which shows Status of Customers (Active/Inactive) for this we have PDF forms which hav texfields to enter the status of customers.I have developed class that will read these textfields but now iI want that these data which is nothing but the Custome metadata for Customer should get extracted (and not maually entered) while adding this file in Alfresco.&lt;/SPAN&gt;&lt;BR /&gt;&lt;BR /&gt;&lt;SPAN&gt;Can anyone help me out in this?&lt;/SPAN&gt;&lt;BR /&gt;&lt;BR /&gt;&lt;SPAN&gt;I already put up two questions on forum about extracting metadata but iI think nobody knows how to do it or nobody is interested to do so.&lt;/SPAN&gt;&lt;/BODY&gt;&lt;/HTML&gt;</description>
      <pubDate>Mon, 07 May 2007 05:00:20 GMT</pubDate>
      <guid>https://connect.hyland.com/t5/alfresco-archive/indexing-of-pdf-custom-fields/m-p/40010#M21352</guid>
      <dc:creator>amarendrakt</dc:creator>
      <dc:date>2007-05-07T05:00:20Z</dc:date>
    </item>
  </channel>
</rss>

