pdfminer.six

Commit Graph

Author	SHA1	Message	Date
Pieter Marsman	91d89af788	Add section to documentation with howto for image extraction (#427 ) * Make structure of documentation more clear: tutorials, how-to, topics and reference * Add howto for images * Restructure tutorials section, and add install section * Always use up-to-date version * Fix indentation warning in docstring * Add option to dumppdf.py and pdf2txt.py to show version Fixes #162	2020-05-17 17:48:06 +02:00
fabbox	7eff108fa5	add shebang line to script in tools (#408 ) * add shebang line to script in tools * fix: use shebang line with python 3 * Moved changelog to unreleased Co-authored-by: Pieter Marsman <pietermarsman@gmail.com>	2020-04-28 10:58:42 +02:00
Jake Stockwin	e55560f858	Fix #395 : Update documentation for boxes_flow, allow None (#396 ) * Update documentation for boxes_flow, allow None * Apply comments from code review * Small wording changes, remove unnecessary comment * Update boxes_flow documentation for pdf2text * Pin version of tox to ensure python 3.4 support	2020-03-26 23:03:49 +01:00
Jake Stockwin	518b5d6efc	Fix #390 : Updated misleading documentation about word_margin (#407 ) * Updated misleading documentation about word_margin * Small change in sentence about word_margin * Remove confusing sentence about adding spaces Co-authored-by: Pieter Marsman <pietermarsman@gmail.com>	2020-03-26 23:02:48 +01:00
Pieter Marsman	1d773dc38a	Fix grouping textlines when bounding box of parent container is wrong (#386 ) * Default value for --all-texts should be false, because using the flag enables it * Fix edge case: when no neighbors are found a line should form its own text box * Added test for grouping textlines where 1 is outside the parent bounding box * Added CHANGELOG.md line	2020-03-14 10:33:39 +01:00
Pieter Marsman	3502dc9f3b	Drop support for legacy Python 2 (#346 ) * Drop support for legacy Python 2 * Add python_requires to help pip * Upgrade Python syntax with pyupgrade * Upgrade Python syntax with pyupgrade --py3-plus * Python 3 imports * Replace six * Update CONTRIBUTING.md * Added line to changelog Co-authored-by: Hugo van Kemenade <hugovk@users.noreply.github.com>	2020-01-04 16:47:07 +01:00
Pieter Marsman	f3ab1bc61e	Enforce pep8 coding-style (#345 ) * Code Refractor: Use code-style enforcement #312 * Add flake8 to travis-ci * Remove python 2 3 comment on six library. 891 errors > 870 errors. * Remove class and functions comments that consist of just the name. 870 errors > 855 errors. * Fix flake8 errors in pdftypes.py. 855 errors > 833 errors. * Moving flake8 testing from .travis.yml to tox.ini to ensure local testing before commiting * Cleanup pdfinterp.py and add documentation from PDF Reference * Cleanup pdfpage.py * Cleanup pdffont.py * Clean psparser.py * Cleanup high_level.py * Cleanup layout.py * Cleanup pdfparser.py * Cleanup pdfcolor.py * Cleanup rijndael.py * Cleanup converter.py * Rename klass to cls if it is the class variable, to be more consistent with standard practice * Cleanup cmap.py * Cleanup pdfdevice.py * flake8 ignore fontmetrics.py * Cleanup test_pdfminer_psparser.py * Fix flake8 in pdfdocument.py; 339 errors to go * Fix flake8 utils.py; 326 errors togo * pep8 correction for few files in /tools/ 328 > 160 to go (#342) * pep8 correction for few files in /tools/ 328 > 160 to go * pep8 correction: 160 > 5 to go * Fix ascii85.py errors * Fix error in getting index from target that does not exists * Remove commented print lines * Fix flake8 error in pdfinterp.py * Fix python2 specific error by removing argument from print statement * Ignore invalid python2 syntax * Update contributing.md * Added changelog * Remove unused import Co-authored-by: Fakabbir Amin <f4amin@gmail.com>	2019-12-29 21:20:20 +01:00
Martin Hasoň	78f06225b6	Removed duplicated and therefore unused code from pdf2txt.py (#341 )	2019-12-09 22:04:05 +01:00
Pieter Marsman	bc034c8e59	Create sphinx documentation for Read the Docs (#329 ) Fixes #171 Fixes #199 Fixes #118 Fixes #178 Added: tests for building documentation and example code in documentation Added: docstrings for common used functions and classes Removed: old documentation	2019-11-07 21:12:34 +01:00
Martin Hasoň	ed1b09c6f2	Fix debug logging for pdf2txt.py and dumppdf.py (#325 ) Fixes #313	2019-11-06 21:47:19 +01:00
Pieter Marsman	33b16b3f07	Deprecate the use of _py2_no_more_posargs (#328 ) Fixes #324	2019-11-02 10:29:39 +01:00
Wm Bentley	495c92e050	Move argparse object setup out of main to separate function. As preparation for implementing Sphinx documentation, create a separate function that builds and returns the argparse parser. Move import argparse out of main to the top of the file.	2018-08-12 21:07:52 -07:00
Andy Kluger	ed7d8308d9	-P is not for page numbers, but passwords, so reflect that in the help text	2018-04-03 12:26:01 -04:00
Antonio Ercole De Luca	0fdebc6739	Removing all the "#!/usr/bin/env python" lines, they do not need for … (#34 ) * Removing all the "#!/usr/bin/env python" lines, they do not need for python3, solving issue number: #19. * Restored all the shebangs in the tools and tests folders (because they are real executables) but used "#!/usr/bin/env python" instead of "#!/usr/bin/python" as this blog points out: https://www.peterbe.com/plog/importance-of-env Removed also the shebang from pdfminer/psparser.py file.	2016-11-08 20:01:11 +01:00
Ivan Teoh	2c8f226907	Fix issues #20 - NameError: global name 'ImageWriter' is not defined	2016-04-26 12:38:42 +10:00
Chris Hager	2e1be5721f	removed settings.ENFORCE_CHECK_EXTRACTABLE	2015-11-01 22:34:18 +01:00
Chris Hager	b686dd0139	pdfminer/settings.py for STRICT and added ENFORCE_CHECK_EXTRACTABLE	2015-11-01 22:28:08 +01:00
Cathal Garvey	268e9fb2bd	Removed typechecking, nothing's exploded yet and argparse does lots of heavy lifting already.	2015-05-30 17:05:28 +01:00
Cathal Garvey	b3553cef10	Cleaning up pdf2txt.py after the partition/move.	2015-05-30 17:03:55 +01:00
Cathal Garvey	cbe270a4bf	Killed the old main function for pdf2txt.py	2015-05-30 16:37:22 +01:00
Cathal Garvey	ead8e778a6	Successfully compartmentalised code, getting closer to moving pdf->text as a module function.	2015-05-30 16:27:58 +01:00
Cathal Garvey	08cb217983	Progress, progress.. not nearly atomic enough, sorry.	2015-05-30 16:14:24 +01:00
Cathal Garvey	1b47bed306	Many changes to make pdf2txt.py work better in Py3, some in that script, others in module! Sorry, changes should have been more atomic. In pdf2txt.py: * Re-wrote main function to use argparse instead of optparse. * Manually tested in Py2/Py3 to get partial consistency. * Errors abound including Tags mode, but most modes weren't working at all in Py3 anyway. * Py2 mode probably unchanged, cannot find any bugs yet... * Kept old main function for posterity, for now. In utils: * Added a few compatibility functions (some string hax required chardet, new dependency): - make_compat_bytes(in_str)-> (py3->bytes \| py2->str) - make_compat_str(in_str)-> (str) - compatible_encode_method(bytesorstring, encoding, erraction)-> (str) In pdfdevice: * To handle different output filetypes in Py3, injected lots of calls to new utils methods, as well as some six.PYX checks and logic. These changes are largely responsible for enhanced Py2/Py3 consistency. In converter: * To handle output filetypes in Py2, injected a few checks and fixes particularly around the py2 `str.encode` method and its assumed usual use-analogies in Py3.	2015-05-17 21:08:57 +01:00
cybjit	2639b15ef4	guess argv encoding in py2 using sys.stdin.encoding	2014-09-16 23:17:26 +02:00
cybjit	14585987c3	keep password api unicode, latin1 or utf-8 is encoded in handler	2014-09-16 22:58:25 +02:00
cybjit	714423883c	setup logging for pdf2txt and fix dumppdf	2014-09-12 00:29:31 +02:00
cybjit	0a2d90c051	pdf2txt: do not double encode stdout	2014-09-07 18:34:11 +02:00
unknown	29c07ea770	Python 3.4 support and tests	2014-09-03 15:26:08 +02:00
Yusuke Shinyama	44074b42ea	Added: stripcontrol for XMLConverter (-S option)	2014-06-22 00:33:00 +09:00
Yusuke Shinyama	1384a3fe8d	Code cleanup: removed some debug flags.	2014-06-14 15:43:10 +09:00
Yusuke Shinyama	bb6f9b6fc9	Added: -R option.	2013-11-25 18:21:19 +09:00
Yusuke Shinyama	d3730a29ec	API change: process_pdf -> PDFPage.get_pages	2013-10-22 18:59:16 +09:00
Yusuke Shinyama	0ea08890d4	renamed: python2 -> python.	2013-10-17 23:05:27 +09:00
Yusuke Shinyama	2221163b94	Split pdfparser.py and pdfdocument.py.	2013-10-10 18:29:30 +09:00
Yusuke Shinyama	82ff98c7b3	imagewriter now works with text output	2011-11-07 01:15:10 +10:00
Yusuke Shinyama	dc8fde0e47	added CCITTFaxFilter support and a very crude image extraction.	2011-07-18 21:07:00 +10:00
Yusuke Shinyama	fcf0d74ecc	tweaks for debugging	2011-04-21 22:07:52 +09:00
Yusuke Shinyama	4918d59bc2	disable caching support	2011-03-03 00:04:43 +09:00
Yusuke Shinyama	7dbb664db3	code cleanup and more debugging options	2011-02-14 23:42:05 +09:00
Yusuke Shinyama	cbd58121e3	fix aggressive vertical writing detection (which ruins layout)	2011-02-02 23:09:34 +09:00
Yusuke Shinyama	d3bcc0eef5	another minor fix	2010-12-26 19:30:46 +09:00
Yusuke Shinyama	a24c452ba2	boxes_flow patch by Daniel Gerber	2010-12-26 17:26:39 +09:00
yusuke.shinyama.dummy	2bf9c23801	check_extractable paramater added git-svn-id: https://pdfminerr.googlecode.com/svn/trunk/pdfminer@276 1aa58f4a-7d42-0410-adbc-911cccaed67c	2010-11-23 10:53:28 +00:00
yusuke.shinyama.dummy	7374b81383	htmlconverter improved git-svn-id: https://pdfminerr.googlecode.com/svn/trunk/pdfminer@274 1aa58f4a-7d42-0410-adbc-911cccaed67c	2010-11-14 15:04:28 +00:00
yusuke.shinyama.dummy	509ab66319	stay with python2 git-svn-id: https://pdfminerr.googlecode.com/svn/trunk/pdfminer@264 1aa58f4a-7d42-0410-adbc-911cccaed67c	2010-10-19 09:57:01 +00:00
yusuke.shinyama.dummy	eb535d4106	change PDFPageAggregator -> PDFLayoutAnalyzer git-svn-id: https://pdfminerr.googlecode.com/svn/trunk/pdfminer@213 1aa58f4a-7d42-0410-adbc-911cccaed67c	2010-04-24 13:31:21 +00:00
yusuke.shinyama.dummy	97848409e5	fix xobject resources bug, thanks to Jose Maria git-svn-id: https://pdfminerr.googlecode.com/svn/trunk/pdfminer@209 1aa58f4a-7d42-0410-adbc-911cccaed67c	2010-04-24 04:32:03 +00:00
yusuke.shinyama.dummy	e77a6ba997	-A (all_texts) option added for layout analysis git-svn-id: https://pdfminerr.googlecode.com/svn/trunk/pdfminer@205 1aa58f4a-7d42-0410-adbc-911cccaed67c	2010-04-10 11:30:03 +00:00
yusuke.shinyama.dummy	2e5b92c18a	writing mode detection git-svn-id: https://pdfminerr.googlecode.com/svn/trunk/pdfminer@196 1aa58f4a-7d42-0410-adbc-911cccaed67c	2010-03-25 11:38:47 +00:00
yusuke.shinyama.dummy	ee34d8d549	bugfix (thanks to Brian Berry). Remaining TODOs: automatic testing for vertical texts. Various layout analysis tuning. git-svn-id: https://pdfminerr.googlecode.com/svn/trunk/pdfminer@193 1aa58f4a-7d42-0410-adbc-911cccaed67c	2010-03-22 08:36:39 +00:00

1 2

77 Commits (99f0c09869370d74eb6a27234284b894c33414f4)