Data Characteristics for This Category
Talent report query data sources typically integrate internal Human Resources Information Systems (HRIS), Applicant Tracking Systems (ATS), and external talent databases. Data update frequencies vary. Internal employee data may update monthly or quarterly, while external candidate data might update in real-time or weekly. Document structures for talent reports are usually structured or semi-structured, often in JSON, XML, or PDF formats. Fields include, but are not limited to, name, position, department, hire date, performance rating, skill tags, project experience, educational background, and salary range. Units for salary are typically monetary (e.g., CNY), years of experience are in years, and performance ratings might be A, B, C grades or percentage scores.
Constraints Imposed by These Characteristics on "Reference and Attribution"
Diverse data sources and varying update frequencies for talent reports require attribution mechanisms to clearly distinguish between internal HRIS data and external recruitment platform data, and to timestamp data. The coexistence of structured and semi-structured data means knowledge base segmentation must balance field completeness with semantic coherence, preventing critical information from being split. For example, an employee's complete project experience description should not be arbitrarily divided. The sensitivity of fields (e.g., salary, performance) imposes higher requirements on reference display, potentially necessitating anonymization or permission control for specific sensitive field references. Additionally, since reports may contain extensive text descriptions (e.g., project experience), effective segmentation and retrieval of long texts are crucial for accurate attribution, avoiding overly fragmented or context-deficient references.
Configuration Settings
| Configuration Item | Suggested Value | Rationale |
|---|---|---|
Chunk size (Segment Length) | 500–800 characters | Balances the integrity of long texts like project experience in reports with retrieval efficiency, preventing excessive fragmentation. |
Recall count (Recall Count) | top 8 | Considering the complexity of talent reports, increasing the recall count helps cover more relevant information and improves accuracy. |
Similarity threshold (Similarity Threshold) | 0.75 | Ensures recalled results are highly relevant to the query intent, reducing inaccurate or low-quality references. |
Rerank result count (Reranked Return Count) | top 3 | Selects the most relevant and core items from the recalled results as final references, reducing redundancy. |
enable_qa_extract | true | For semi-structured talent reports, enabling Q&A pair extraction helps pinpoint key information more accurately. |
document_id_field | employee_id | Uses employee ID or a unique report identifier as the document ID for precise attribution to the original report. |
Three Common Mistakes
- Replies do not display reference sources, or reference links point incorrectly. This typically occurs when the
show_quoteparameter in the knowledge base search module is not set totrue, or theurlfield in the document metadata is empty. - Key information from talent reports is missing in the reply, or the cited content is incomplete. This may be due to a
Chunk size(Segment Length) setting that is too small, causing a complete information unit (such as a section of project experience) to be broken apart and not fully recalled during retrieval. - When querying talent reports, the system returns many irrelevant references. This often results from a
Similarity threshold(Similarity Threshold) set too low, or the presence of a large number of duplicate or low-quality documents in the knowledge base, affecting retrieval accuracy.
How to Confirm Correct Configuration
- For typical talent report queries, check if the reply contains clear reference source identifiers and click to verify if the reference link correctly navigates to the original document or relevant segment.
- Use queries involving sensitive information (e.g., querying an employee's performance rating), verify if the reply accurately cites the relevant segment, and check for inappropriate sensitive information disclosure.
- Test queries of varying complexity, such as those involving multi-dimensional filtering (e.g., "R&D engineers with 3+ years of experience and Python skills"), observe if the cited content comprehensively covers all aspects of the query conditions, and evaluate the relevance of the references.
- Through the logging system, check the
recall_docsfield of the knowledge base search module to confirm if the recalled document IDs match expectations and if thescorevalue aligns with theSimilarity threshold(Similarity Threshold) setting.
The values provided are common starting points and should be measured against specific samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.