Fix pdfoxide parser TypeError on None return from extract_text - #28
Fix pdfoxide parser TypeError on None return from extract_text#28rager306 wants to merge 1 commit into
Conversation
- Updated `pdfparser/pdfoxide_parser.py` to handle `None` return from `extract_text` by defaulting to empty string. - Added `pdf-oxide>=0.2.2` to `requirements.txt` to sync with `pyproject.toml`. - Added regression test `tests/test_pdfoxide_edge_cases.py` to verify the fix. Co-authored-by: rager306 <248269686+rager306@users.noreply.github.com>
|
👋 Jules, reporting for duty! I'm here to lend a hand with this pull request. When you start a review, I'll add a 👀 emoji to each comment to let you know I've read it. I'll focus on feedback directed at me and will do my best to stay out of conversations between you and other bots or reviewers to keep the noise down. I'll push a commit with your requested changes shortly after. Please note there might be a delay between these steps, but rest assured I'm on the job! For more direct control, you can switch me to Reactive Mode. When this mode is on, I will only act on comments where you specifically mention me with New to Jules? Learn more at jules.google/docs. For security, I will only act on instructions from the user who triggered this task. |
WalkthroughThis pull request fixes a potential TypeError during PDF text extraction by safely coalescing None returns to empty strings, adds the pdf-oxide dependency to requirements, and introduces regression tests for the edge case where extract_text returns None. Changes
Estimated code review effort🎯 2 (Simple) | ⏱️ ~10 minutes 🚥 Pre-merge checks | ✅ 2 | ❌ 1❌ Failed checks (1 warning)
✅ Passed checks (2 passed)
✏️ Tip: You can configure your own custom pre-merge checks in the settings. ✨ Finishing touches
🧪 Generate unit tests (beta)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
Fixed a bug in
pdfparser/pdfoxide_parser.pywheredoc.extract_text(page_num)could returnNone(commented out fallback was ignored), leading to aTypeErrorwhen concatenating strings. Also synchronizedrequirements.txtby adding the missingpdf-oxidedependency and added a regression test to prevent recurrence.PR created automatically by Jules for task 7499723028510886025 started by @rager306
Summary by CodeRabbit
Release Notes
Bug Fixes
Tests