TextIn xParse Turns a Mess of Supplier Quotes into a Traceable Procurement Report
Recently, a procurement colleague mentioned a very specific task: comparing several sensor suppliers for a bridge and tunnel structural health monitoring project. The procurement team cares not only about the total price but also needs to verify the parameters and quantities of fiber optic strain, crack displacement, inclination, and vibration sensors, as well as the channel counts of data acquisition units and edge gateways, delivery cycles, warranty periods, and maintenance conditions. The trouble is that Supplier A sends a multi-level header Excel, Supplier B sends a Word document with 'tables within tables,' Supplier C sends a scanned PDF, and Supplier D uses a quotation sheet with package-level discounts; the package-level total price, equipment sub-items, optional accessories, and service terms are often not on the same layer.
I frequently use WorkBuddy and have previously used TextIn to build invoice and contract assistants. After seeing the TextIn xParse connector in WorkBuddy, I used this set of desensitized simulation materials for a verification: without manually organizing the files first, I directly fed the procurement requirements and the four original supplier quotations to it to see if it could ultimately generate a supplier comprehensive evaluation report ready for further review.
First, Look at the Final Delivered Result
This time, a total of 6 files were used: the procurement team's requirement Word document and point location Excel, as well as the Excel, Word, and PDF files submitted by Suppliers A, B, C, and D. After parsing each file with TextIn xParse, the four suppliers' effective tax-inclusive quotations, B-01 frequency, delivery cycles, warranty periods, and risk levels were placed into a single comprehensive evaluation table. Additionally, 7 cross-referencing anomalies and caliber issues were listed.
This result did not directly rank the lowest price first. Supplier B had the lowest quotation, but the 0.8Hz low-frequency lower limit of B-01 constituted a hard negative deviation; Supplier D had a mid-range price, a 36-month warranty, and included a package-level discount, making it the conditional first candidate; Supplier A had the shortest delivery cycle but needed to supplement the B-01 frequency proof. Uncertain quantities, amounts, and parameters were not automatically filled in but were left in the manual review checklist.
If you just want a quick initial screening and don't want to read the full report first, you can also directly ask it to output a comparison table based on the dimensions you care about. I asked another question, placing effective quotation, mandatory equipment completeness, B-01 frequency, D-01 channel count, technical deviations, optional accessories, package-level discounts, delivery, warranty, service response, installation and commissioning, training and maintenance, quotation cross-referencing, document completeness, risk basis, and recommendation conclusions into the same table.
This kind of multi-dimensional table is more suitable for ad-hoc discussions and quick screening, where key differences can be found at a glance. When I need to keep the original evidence, cross-referencing anomalies, and recommendation basis intact, or when entering the subsequent review process, I ask it to generate a detailed report.
Finally, an 18-page, 5955-word Word version of the 'Supplier Comprehensive Evaluation Report' was obtained. The supplier summary table, quotation caliber, technical deviations, business risks, recommendation order, and items to be confirmed are all inside, allowing procurement personnel to continue editing, circulating, and archiving.
The final deliverable is this report, and the preceding document parsing is the foundation. Let's start with the connection of TextIn xParse and two rounds of actual testing, and finally organize the verified procurement rules into a directly reusable Agent.
First, Connect the Connector
Searching for TextIn xParse on WorkBuddy's connector page, the entry point is not hard to find. It doesn't simply convert files into a block of plain text; instead, it first calls a parsing service to organize the text, tables, and layout relationships in Excel, Word, and PDF files, then hands them over to WorkBuddy for cross-file analysis. The connector page shows currently available file types like PDF, Excel, Word, which perfectly matches the mixed state of procurement materials.
A set of authorization codes will pop up during authorization. Account, key, and full token are all sensitive information; reusable real credentials cannot be placed in the article. The authorization code is only used for the current connection process. After entering it, return to WorkBuddy to confirm the authorization object. This step only establishes the connection and does not start parsing files; the actual service call happens later when uploading materials.
After copying the authorization code, return to the authorization page to confirm WorkBuddy. The page will display the connector being authorized and the authorization status. I did not upload files immediately after authorization was complete but first checked the connector status and available capabilities, confirming that Excel, Word, and PDF were all within the support scope, also avoiding mistaking the success prompt on the authorization page for the connector being ready.
After confirming on the authorization page, WorkBuddy's connector status will change to available. Here, I conveniently checked the tool list, supported file types, and call entry points to ensure that when doing the normal parsing comparison later, the only test variable was whether xParse was enabled. If xParse does not appear in the tool list, the subsequent file analysis cannot be considered successfully connected.
The connector popup also shows 1,000 free pages per day. This is enough for testing these few simulation materials, but this is only the quota displayed for the current account. I treat it as the available quota for this verification and do not infer the cost of formal projects from this; actual usage still needs to be evaluated based on the volume of procurement documents.
After the connection was complete, I didn't immediately put in all four suppliers' complex materials. Instead, I first ran a baseline with a regular Excel. Simple tables are inherently easy to read, making them more suitable for first confirming whether the speed and basic structure are stable.
First, Run a Test with a Regular Excel
The first test used a previous itemized quotation Excel, containing 4 worksheets: quotation overview, parent-child level itemized quotation, technical parameters, and business deviations. The table wasn't particularly tricky, making it perfect for a baseline: the same file, the same question, only changing whether TextIn xParse was used, to first see the difference in response time and basic results.
This time, I explicitly wrote in the question 'Use TextIn xParse to analyze the content of this file.' WorkBuddy's complete answer took 2 minutes and 08 seconds. The result also marked xParse parsing time as 328ms, and correctly identified the file name, supplier, project, and 4 worksheets. The two times are not the same thing: 328ms is the time for the connector to complete file parsing, while 2 minutes and 08 seconds also includes the time for WorkBuddy to subsequently read, organize, and generate the answer.
The core quotation page was not compressed into a piece of scattered text. xParse restored the parent-child relationship of 5 major systems and 12 equipment and service items, and retained equipment codes, models, quantities, units, tax-exclusive unit prices, tax-inclusive amounts, and delivery batches. Blank rows in the original table that omitted repeated parent-level names could also continue to be categorized under the corresponding A, B, C, D, E system classes.
Recognition didn't stop at the price table. Delivery, warranty, technical deviations, and service risks in the business response page were also organized: total quotation ¥1,309,184.80, delivery in two batches, 30-month warranty; B-01 frequency lower limit better than procurement requirement, SV-01 listed as the only medium-risk item due to different on-site arrival times in different regions. At the bottom of the page, you can also see the Markdown and JSON files generated by xParse. The original Markdown on the right retains merged header information like rowspan and colspan.
Subsequently, I turned off xParse and re-identified the same Excel with the same question. Normal parsing took 3 minutes and 39 seconds. Both times could read the main content of this well-structured table, but comparing the complete answer time, after connecting xParse, it dropped from 219 seconds to 128 seconds, saving 91 seconds of waiting time, with a response speed about 1.7 times that of normal parsing. This is a single actual record for one file, not a large-scale performance test.
The regular Excel mainly verified speed and basic structure. After connecting xParse, the complete answer was faster, and Markdown and JSON preserving table relationships were also obtained. The more obvious differences need to be seen in nested Word documents, scanned PDFs, and cross-file price comparisons.
Put the Four Suppliers' Original Files on the Table
I first fixed the procurement team's requirements: which equipment is mandatory, how point locations convert to quantities, what the maximum delivery cycle is, and what hard thresholds exist for warranty and fault response. The supplier files were kept as received, without pre-converting them into a unified template or manually copying tables first.
Supplier A quoted using Excel. Its quotation page simultaneously presented procurement packages, equipment bodies, supporting accessories, implementation services, package subtotals, full tax-inclusive totals, and mandatory tax-inclusive totals, with merged cells in multi-level headers. If procurement personnel only copy the last total price, they can easily include optional accessories or package-level discounts.
Supplier B submitted a Word technical response. The outer table first responded by procurement package, while the inner layer embedded a configuration package detail, followed by itemized quotations and business conditions. Equipment parameters, quotations, and proof documents were not in the same table; whether a certain model had corresponding accessories needed to be judged by piecing together several locations.
Supplier C's PDF was closer to the scanned copies actually received by procurement personnel: complex headers, table boundaries not always clear, some information scattered across different pages, and finally revision pages and handwritten supplements. For such files, just reading out the characters is not the end; table hierarchies, page relationships, and revision priorities must also be preserved.
Supplier D's quotation first used Excel as the data source, then exported it as a watermarked PDF. The quantity for A-02 crack displacement sensor was deliberately changed to 2, while the unit price and amount kept the original values, to check whether the parsing result could list this cross-referencing anomaly for review.
After the files were laid out, the question was no longer 'Can it recognize Chinese characters?' but whether it could map different suppliers' writing styles to the procurement team's codes like A-01, A-02, B-01, D-01, E-02, and retain the evidence source. Next, let's look at the results without and with xParse side by side.
Effect Comparison: With or Without xParse, What's the Difference
Without xParse, What Does Normal Parsing Miss
The first joint test deliberately did not use TextIn xParse. I uploaded the procurement Word, procurement Excel, Supplier A Excel, Supplier B Word, and Supplier C PDF, using the Deepseek-V4-Flash model. The question was also phrased in the procurement personnel's tone: compare sensors, acquisition equipment, tax-inclusive prices, delivery, warranty, and deviation items, and finally point out areas needing manual review.
Normal parsing is not completely unusable. It could find conflicts in the procurement team's quantity caliber and catch keywords like Supplier B's B-01 frequency lower limit and Supplier C's quotation revision page. However, on the Word technical response, the model judged PK-D and PK-E as 'missing quotations', while the original file actually placed the quotation in another layer of the table. The text was read, but the relationships between tables were not preserved.
To confirm where the problem lay, I took out Supplier B's Word separately and used the same Hy3 model for comparison. Without xParse, the result split the outer configuration package and itemized quotation into two unrelated areas, and PK-D and PK-E could not be matched.
Normal models can quickly scan text and find keywords, but procurement price comparison requires traceable fields. Missing one nested table could misjudge an entire package of equipment as a missing item. The problem was clear; next, keeping the files and model unchanged, only turn on TextIn xParse.
After Connecting xParse, Table Relationships Returned
After enabling xParse and asking the same question, the results included information that normal parsing did not stably retain: original file names, page numbers or worksheet names, hierarchies formed by merged headers, relationships between embedded tables and outer procurement packages, and the evidence location for each conclusion. WorkBuddy handled the cross-file comparison, while xParse first restored the file structure.
In the follow-up results, Supplier B's official name was completed, Supplier A's three-layer header and evidence location were preserved, and the point location split in the procurement Word could also correspond to equipment quantities. On the right side, multiple Markdown files output by xParse could be seen, with tables preserved in HTML table form, providing stable input for subsequent normalization.
The same Hy3 model continued to analyze Supplier B's Word, and the conclusions could then be verified: PK-A and PK-B lacked embedded configurations; PK-C's outer response was missing but the quotation had C-01 and C-02; PK-D and PK-E's outer, nested, and quotation parts could correspond; B-01's response range was 0.8-120Hz, lower than the procurement requirement's low-frequency lower limit; attachments were not embedded and needed the supplier to supplement proof.
I also compared the mandatory equipment and services of the three suppliers according to the procurement execution caliber. A-01 quantity conflict, B-01 frequency, D-01 channel count, and E-02 warranty period were listed separately, and amounts were calculated based on mandatory items, without directly stuffing optional accessories into the lowest effective quotation.
In the normalization results for the early three-sample set, Supplier A's mandatory tax-inclusive amount was about 1.3227 million yuan, Supplier B's about 1.2671 million yuan, and Supplier C's about 1.3486 million yuan based on the revision page caliber. These numbers only apply to the current procurement list, tax rate, and mandatory definition. The role of xParse is to allow amounts, technical deviations, and missing evidence to be traced back to the original rows, rather than just leaving an untraceable price ranking.
At this point, the speed for regular Excel, structural restoration for complex materials, and cross-file price comparison had all been verified. A single conversation could already complete the task, so I further organized this set of procurement calibers into a reusable Agent.
Finally, Turn This Process into a Reusable Agent
I wrote the role, parsing requirements, normalization rules, and report structure into the Agent's System Prompt. In the future, when receiving new supplier materials, just replace the uploaded files without needing to re-explain how to judge mandatory, optional, discounts, and technical deviations. Unrecognizable fields are marked as 'Unrecognized,' and fields not written in the original file are marked as 'Not Provided'; no guessing is allowed to fill in the table.
This System Prompt first requires TextIn xParse to restore the document structure, then unify the procurement caliber according to equipment codes, and finally output an evaluation report traceable to page numbers or worksheets. The complete original text can be directly reused:
[Role]
You are a procurement decision assistant for a bridge and tunnel structural health monitoring project, responsible for initial screening, normalized price comparison, technical deviation checking, and procurement risk identification of supplier materials.
Your task is to provide evidence-based decision support for procurement personnel. You must not replace procurement personnel in making final approvals, and you must not speculate on conclusions when data is missing.
[Task Objective]
Based on the procurement requirements, point location list, review rules, and the Excel, Word, PDF materials submitted by each supplier, generate a 'Supplier Comprehensive Evaluation Report' and output a Word document that can be further edited and archived.
[Tool Usage Requirements]
1. Must first call TextIn xParse file by file to parse all procurement and supplier documents. Do not rely solely on file names, file summaries, or plain text extraction to judge content.
2. After parsing each file, retain the file name, page number or worksheet name, header hierarchy, merged cells, nested tables, parent-child relationships, and original text.
3. Prioritize using the Markdown text and table structure data returned by xParse for subsequent analysis.
4. If a file fails to parse or its content is incomplete, record the file name, failure reason, and missing scope. Do not skip it and assume it meets requirements.
[Document Parsing Requirements]
1. Identify the hierarchical relationships among procurement packages, subsystems, equipment bodies, supporting accessories, implementation services, and package-level discounts.
2. Extract equipment code, equipment name, manufacturer model, quantity expression, calculated quantity, unit, technical parameters, tax-exclusive amount, tax rate, tax-inclusive amount, delivery cycle, warranty period, and service response time.
3. Distinguish between mandatory items, bundled accessories, and optional accessories; package subtotals, totals, and grand totals must not be double-counted into equipment details.
4. When outer tables, nested configuration tables, and attachment references exist in Word, establish corresponding relationships; do not misjudge nested tables as missing.
5. When scanned tables, handwritten revisions, cross-page content, or revision pages exist in PDF, retain the original value, revised value, and corresponding page number simultaneously.
6. Mark unrecognizable information as 'Unrecognized,' and information not provided in the original file as 'Not Provided'; do not fill in on your own.
[Procurement Normalization Rules]
1. Use the equipment codes and procurement packages in the procurement team's documents as the primary comparison keys to unify different suppliers' fields into the same comparison caliber.
2. Process equipment bodies, mandatory accessories, bundled accessories, optional accessories, implementation services, package-level discounts, package subtotals, and totals separately.
3. The effective tax-inclusive quotation only calculates mandatory equipment, bundled accessories, and mandatory services within the procurement scope; optional accessories are listed separately and must not be mixed into the lowest effective quotation.
4. When the same code appears with multiple quantities, models, parameters, or amounts in different files, do not overwrite directly; retain all original values and mark the source of the conflict.
5. Quantity, unit price, tax-exclusive amount, tax amount, and tax-inclusive amount need cross-referencing verification; list items that cannot be cross-referenced for manual review.
6. When the supplier's self-reported mandatory scope differs from the procurement team's scope, retain both calibers separately and explain the impact on the effective quotation.
[Review Requirements]
1. Technical parameters, delivery, warranty, service response, and implementation scope must be judged based on the procurement team's requirements.
2. Items clearly not meeting hard conditions are marked as 'Negative Deviation'; items with missing evidence or content conflicts are marked as 'To Be Confirmed'.
3. Do not recommend suppliers based solely on the lowest price. Recommendation reasons must simultaneously consider technical satisfaction, effective quotation, delivery, warranty, implementation completeness, and risk.
4. Each risk, deviation, and anomaly must be annotated with the supplier, equipment code, original value, judgment basis, and file name, page number, or worksheet location.
5. All recommendations are conditional initial screening suggestions; separately list manual confirmation items that could change the recommendation order.
[Report Output]
Generate a 'Supplier Comprehensive Evaluation Report,' containing at least:
1. List of parsed files and description of TextIn xParse parsing results.
2. Supplier comprehensive evaluation table: supplier, effective tax-inclusive quotation, technical match, delivery cycle, warranty period, risk level, and preliminary suggestion.
3. Completeness comparison of mandatory equipment, bundled accessories, and implementation services.
4. Optional accessories, package-level discounts, and quotation caliber differences.
5. List of technical deviations and business risks.
6. Cross-referencing anomalies for quantity, unit price, tax amount, tax-inclusive amount, and cross-page content.
7. Supplier recommendation order, recommendation reasons, and reasons for not recommending as the first choice.
8. Items that must be confirmed by procurement personnel, with file name, page number, or worksheet location attached.
Complete all file parsing and cross-verification first, then output conclusions. Report conclusions must be traceable back to the original materials; do not output deterministic judgments without evidence locations.
During actual execution, all 6 files generated Markdown text and table structure data, totaling about 1.3MB. Page numbers, worksheet names, merged cells, nested tables, and parent-child hierarchies were still retained, allowing WorkBuddy to perform cross-file normalization by equipment code later.
Previously, when receiving supplier materials, one usually had to first expand merged cells, fill in parent levels, split Word embedded tables, and copy PDF revision pages back into Excel before starting price comparison. Now, when suppliers send materials in different formats, they can first be handed to this Agent: TextIn xParse retains text, tables, hierarchies, and evidence locations, and WorkBuddy then organizes them into the same comparison caliber according to procurement rules.
In the comparison of simple Excel files, TextIn xParse's complete response was faster; when switching to multi-level header Excel, nested Word, and scanned PDFs, the structure it retained allowed quotations, parameters, deviations, and anomalies to continue being verified. What is finally obtained is not just a few paragraphs of parsed text, but a supplier comprehensive evaluation Word report that can continue to be edited, circulated, and archived.
This report is used for procurement initial screening and decision support. Quotation validity periods, payment terms, original manufacturer authorizations, missing attachments, and technical deviations still need to be confirmed by procurement personnel, and final approval will not be handed over to the tool. TextIn xParse first solidly handles the most time-consuming document organization and evidence location, so procurement personnel can focus their energy on areas that truly require judgment.
Finally, a reminder again on how to add TextIn xParse in Workbuddy: