用 Python 提取 PDF 文本的简单方法（用我的手指搅乱吧樱花未增删翻译无马）-eolink官网

用 Python 提取 PDF 文本的简单方法（用我的手指搅乱吧樱花未增删翻译无马）

你好，我是征哥，一般情况下，Ctrl+C 是最简单的方法，当无法 Ctrl+C 时，我们借助于 Python，以下是具体步骤：

第一步，安装工具库

1、tika — 用于从各种文件格式中进行文档类型检测和内容提取

2、wand — 基于 ctypes 的简单 ImageMagick 绑定

3、pytesseract — OCR 识别工具

创建一个虚拟环境，安装这些工具

python -m venv venvsource venv/bin/activatepip install tika wand pytesseract

第二步，编写代码

假如 pdf 文件里面既有文字，又有图片，以下代码可以直接识别文字：

import ioimport pytesseractimport sysfrom PIL import Imagefrom tika import parserfrom wand.image import Image as witext_raw = parser.from_file("example.pdf")print(text_raw['content'].strip())

这还不够，我们还需要能失败图片的部分：

def extract_text_image(from_file, lang='deu', image_type='jpeg', resolution=300): print("-- Parsing image", from_file, "--") print("---------------------------------") pdf_file = wi(filename=from_file, resolution=resolution) image = pdf_file.convert(image_type) image_blobs = [] for img in image.sequence: img_page = wi(image=img) image_blobs.append(img_page.make_blob(image_type)) extract = [] for img_blob in image_blobs: image = Image.open(io.BytesIO(img_blob)) text = pytesseract.image_to_string(image, lang=lang) extract.append(text) for item in extract: for line in item.split("\n"): print(line)

合并一下，完整代码如下：

import ioimport sysfrom PIL import Imageimport pytesseractfrom wand.image import Image as wifrom tika import parserdef extract_text_image(from_file, lang='deu', image_type='jpeg', resolution=300): print("-- Parsing image", from_file, "--") print("---------------------------------") pdf_file = wi(filename=from_file, resolution=resolution) image = pdf_file.convert(image_type) for img in image.sequence: img_page = wi(image=img) image = Image.open(io.BytesIO(img_page.make_blob(image_type))) text = pytesseract.image_to_string(image, lang=lang) for part in text.split("\n"): print("{}".format(part))def parse_text(from_file): print("-- Parsing text", from_file, "--") text_raw = parser.from_file(from_file) print("---------------------------------") print(text_raw['content'].strip()) print("---------------------------------")if __name__ == '__main__': parse_text(sys.argv[1]) extract_text_image(sys.argv[1], sys.argv[2])

第三步，执行

假如 example.pdf 是这样的：

在命令行这样执行：

python run.py example.pdf deu | xargs -0 echo > extract.txt

最终 extract.txt 的结果如下：

-- Parsing text example.pdf -----------------------------------Title pure textContent pure text Slide 1 Slide 2----------------------------------- Parsing image example.pdf -----------------------------------Title pure textContent pure textTitle in imageText in image

你可能会问，如果是简体中文，那个 lang 参数传递什么，传 'chi_sim'，其实是有官方说明的，链接如下：

PDF 中提取文本的脚本实现并不复杂，许多库简化了工作并取得了很好的效果，如果你知道从 PDF 或任何文件中提取文本的其他方法，请留言告诉我。

Python接口自动化之文件上传/下载接口怎么实现

513 2022-09-06

用 Python 提取 PDF 文本的简单方法（用我的手指搅乱吧樱花未增删翻译无马）

java中的接口是类吗

Spring中的aware接口详情

Python接口自动化之文件上传/下载接口怎么实现

推荐文章

接口调用是什么意思？几种常用接口调用方式

接口设计原则

8款在线 API 接口文档管理工具

api管理系统是什么？

什么是接口调试？接口调试的步骤有哪些？

api 接口管理系统有哪些？

接口测试有几种测试方法

API文档生成工具有哪些？

微服务和api网关区别

交换机配置步骤

最近发表

热评文章

在线接口文档管理工具推荐，支持在线测试，HTTP接口

开源的在线接口文档wiki工具Mindoc的介绍与使

如何优雅的进行接口设计？接口设计的六大原则是什么？

什么是API测试,api检测公司

软件接口设计怎么做？前后端分离软件接口设计思路

接口管理平台推荐，几大接口管理平台总有一款适合你！