Latest Activity

Profile IconPriyanka Beniwal, Khushi Jain, Meghna S and 10 more joined LIS Links
8 hours ago
Dr MANIKANDAN T posted a blog post
8 hours ago
Dr. A. Madhava Rao posted a blog post
8 hours ago
Gopal Pandey posted an event
8 hours ago
SHEIK MAIDEEN posted an event

Webinar on DELNET: Resources and Services at Online

September 9, 2026 from 3pm to 4pm
8 hours ago
Kapil Sunehria posted an event

UGC-MMTTP Online Two Week Refresher Course in Library and Information Science at RAMANUJAN COLLEGE, UNIVERSITY OF DELHI

August 31, 2026 at 10:30pm to September 14, 2026 at 5:30pm
8 hours ago
OSS Prasad posted an event

72nd ILA Annual Conference – International Conference on “Future Libraries: Advancing Research and Excellence (FLARE 2027)” at Indira Gandhi Memorial Library, University of Hyderabad, Gachibowli

January 7, 2027 to January 9, 2027
8 hours ago
Shompa Das (Chaudhury) posted a discussion
9 hours ago
Dr. Kalpana. T M updated their profile
11 hours ago
sangeeta sharma is attending Sunita Pareek's event
Saturday
Dr MANIKANDAN T updated their profile
Friday
DHARITRI DEORI and vasudev tiwari are now friends
Friday
vasudev tiwari updated their profile
Friday
Satbir Chauhan updated an event
Thumbnail

ETD - 2026 : 29th International Symposium of Electronic Theses and Dissertations (ETDs) at IIT Delhi

October 23, 2026 at 9am to October 25, 2026 at 6pm
Friday
Dr. Bebi updated their profile
Thursday
Santosh Kumar Kori updated their profile
Thursday
DHARITRI DEORI and Ramakrishnan are now friends
Thursday
Dr. Anil Kumar Jharotia shared their photo on Facebook
Wednesday
Rajesh Meshram updated their profile
Wednesday
Ramakrishnan updated their profile
Wednesday

Dear Friends

 

We are in need of a PDF Metadata Extractor Information, preferably free and not online. Please share the information if anybody using it. Actually it is for using in combination with DSpace software, but we can not go online with our collection.

Any help will be highly appreciated.

Thank you

Subeesh A C

Views: 1050

Reply to This

Replies to This Forum

Try ExitTool

http://www.sno.phy.queensu.ca/~phil/exiftool/

I have been using it for extracting metadata from PDFs for using in DSpace.  It is possible to extract metadata from all PDFs at one go, if you are familiar with command line options.

S. Baskar

Thank you very much sir

But I think the tool is extracting data from document properties in my try. Are you getting the appropriate data with exiftool?

Subeesh A C

Hi,

Using the below command, you can extract all metadata (i.e. all metadata tags associated with the PDF document) from hundreds of PDF documents and save it as CSV file which could be used for doing batch import within DSpace.  

In case, if you require only specific tags, then you have to mention the required metadata tags for extracting.  I have given an example below for your understanding.

To extract all available metadata tags from the PDF documents and save it as a CSV file

---------------------------------------------------------------------------------------------------------------------

exiftool -csv  *.pdf > output.csv

To extract specific metadata tags from the PDF documents and save it as a CSV file

-----------------------------------------------------------------------------------------------------------------------------

exiftool  -TAG -Title   -TAG -Author  -TAG -Producer  -TAG -Subject -TAG -Description -TAG -Type -TAG -Keywords -TAG -ISBN -TAG -Isbn -TAG -Createdate -TAG -CourseID  -TAG -FileSize -TAG -PageCount -TAG -PDFVersion -d %Y-%m-%d  *.pdf -csv > output.csv

Hope this helps.


S. Baskar

LinuXpert Systems

ExifTool Tag Names

The tables listed below give the names of all tags recognized by ExifTool.

http://www.sno.phy.queensu.ca/~phil/exiftool/TagNames/index.html

Thank you very much sir

I have created a small uitlity for extracting information from pdf files  few years ago . it will extract data from all files in a folder and save in tab delimited text file.

you can try it. hope it helps. pls let me know.

i have uploaded the program to google drive. Click here to download

with regards

Mujib Rahiman

KV Kanjikode

Thanks sir, I will surely let you know.

Regards 

Subeesh A C

Sir

I have checked your software, its a great effort if you have coded it yourself. As I see most of the software(s) are not able to identify the pdf files metadata as we require. I think the problem is mostly revolve around  the structure of pdf files itself. In my case the pdf files are not having any standard structure (+ OCR ) in it for the algorithm to extract as it did for any appropriate one. Since we are in hurry and we require more metadata for the current work, we are thinking of indexing it and filtering it later through various categories. Anyway thanks for your reply.

Regards 

Subeesh A C

RSS

© 2026   Created by Dr. Badan Barman.   Powered by

Badges  |  Report an Issue  |  Terms of Service

LIS Links whatsApp