How ARXIVIX Uses AI
ARXIVIX turns the files and sources you select into a searchable library. Some processing uses ordinary software rules. Depending on the services enabled, other steps use AI to read document images or find relevant passages.
What happens to a document
- We receive the file or fetch the source you selected. A watcher may check it again on your chosen schedule. Reference mode keeps a link and a short excerpt while fetched material may be processed temporarily for metadata and alerts.
- We extract text. Text-bearing files can often be read directly. Scanned pages and images may need optical character recognition, or OCR. An external OCR provider receives the file or image needed for that work.
- We prepare a search index. We divide retained text into smaller passages. When an embedding service is enabled, it converts each passage into a list of numbers describing patterns relevant to search. These lists are called embeddings or vectors.
- We store the search index in ARXIVIX's database. The index includes passage text, vectors, and references back to the source. It belongs to the project's access-controlled library. It is not stored on your device as part of the normal web service.
- We retrieve relevant material when you search or ask a question. We may process your question with the embedding service and compare it with the indexed passages. An optional ranking provider may receive your question and candidate passages to improve their order.
An embedding is a search representation, not an anonymous substitute for the original text. We protect and delete it as part of your content. Creating embeddings does not train or fine-tune an AI model.
Current processing services
The processing panel below identifies this deployment’s configured services, model identifiers, information processed, and verification status. It updates from the deployed configuration. Depending on that configuration, answers may use source excerpts or an AI model; metadata and summaries may use local rules or an AI model. External processing stays disabled until the provider's no-training controls are verified.
The panel distinguishes steps performed with AI models from steps performed with ordinary software rules. The available processing services can change as ARXIVIX develops; the listed information is updated when those services change.
Private projects and external processing
Private means access is limited through ARXIVIX's project and workspace permissions. It does not mean every processing step takes place on the same server. Private documents and questions may be sent to the processing providers listed above as part of the features you use.
This processing is part of the service, disclosed when you create an account and accept the Terms. The standard beta does not offer a separate switch that guarantees local-only processing. Only submit information you are authorized to have processed in this way.
Your content is not training data
POPVOX does not use your documents, extracted text, questions, answers, or other submitted content to train or fine-tune AI models. We require the providers processing that content for ARXIVIX to operate under terms and settings that prohibit that training.
Where OpenAI's API is enabled, its published policy excludes API inputs and outputs from model training unless the customer opts in. POPVOX's policy is to keep such sharing disabled. OpenAI may retain abuse-monitoring information, ordinarily up to 30 days, with documented exceptions. Training restrictions and retention are separate controls. OpenAI's data controls
Other providers have their own retention rules, shown with their entries above. Authorized provider personnel may review information for security, abuse prevention, or legal compliance under those rules.
Publishing and deletion
An authorized user must confirm before uploaded material becomes public. For watchers, the confirmation explains whether future collected items will be published automatically. When public full text is disabled, public interfaces expose only permitted metadata; they do not return source excerpts, source-derived summaries, or content-based answers.
Public widgets, feeds, and connections to AI tools allow people and software to retrieve the information you publish. Publishing through ARXIVIX does not itself grant permission to train a model on that material. Independent recipients may retain copies, and we cannot control their subsequent conduct or remove copies they already made.
When you delete material, our deletion process covers the original, extracted text, associated vectors, and service-maintained derivatives, subject to the periods and limited exceptions in the Privacy Policy. Active-system deletion is completed within 30 days of a valid request; backup copies expire within 90 days. Public copies held independently are outside that process.
Check important results
OCR can misread a page. Search can miss relevant material. Automated metadata and answers can be incomplete or incorrect. Use the source references to check important details, and do not treat ARXIVIX output as professional advice or as a determination that material is lawful to use or publish.
Questions about processing or your information: legal@popvox.com.
Processing at a glance
External services stay off until their no-training controls are verified.
- Search embeddings
- ARXIVIX server
- Configured
- Relevance ranking
- ARXIVIX server
- Configured
- Scanned documents (OCR)
- ARXIVIX server
- Configured
- Chat answers
- OpenAI
- Configured
- Document metadata
- OpenAI
- Configured
Models, data & verification details
These details update with the service configuration. Existing indexes may use earlier models; temporary failures may use the local fallback described above.
Search embeddings
- Provider / model
- ARXIVIX server · local-hash-v1
- Data processed
- Document passages and search questions
- Status
- Configured
- Region
- ARXIVIX hosting region
- Last verified
- Not yet recorded
Relevance ranking
- Provider / model
- ARXIVIX server · Local rules
- Data processed
- Search questions and candidate passages
- Status
- Configured
- Region
- ARXIVIX hosting region
- Last verified
- Not yet recorded
Scanned documents (OCR)
- Provider / model
- ARXIVIX server · Tesseract / Poppler
- Data processed
- Document files and images
- Status
- Configured
- Region
- ARXIVIX hosting region
- Last verified
- Not yet recorded
Chat answers
- Provider / model
- OpenAI · gpt-4.1-mini-2025-04-14
- Data processed
- Source passages, questions, and conversation history
- Status
- Configured
- Region
- Default OpenAI API processing (no regional residency configured)
- Last verified
- 2026-09-11
Document metadata
- Provider / model
- OpenAI · gpt-4.1-mini-2025-04-14
- Data processed
- Document text and filenames
- Status
- Configured
- Region
- Default OpenAI API processing (no regional residency configured)
- Last verified
- 2026-09-11