home/categories/llm-ai/patricio0312rev-skills-ai-engineering-doc-to-vector-dataset-generator-skill-md
llm-aidata-ai
doc-to-vector-dataset-generator
Converts documents into clean, chunked datasets suitable for embeddings and vector search. Produces chunked JSONL files with metadata, deduplication logic, and quality checks. Use when preparing "training data", "vector datasets", "document processing", or "embedding data".
maintainer
patricio0312rev
Обновлено 1/12/2026
Звёзды
6
Форки
0
quick start
Installation and usage
Converts documents into clean, chunked datasets suitable for embeddings and vector search. Produces chunked JSONL files with metadata, deduplication logic, and quality checks. Use when preparing "training data", "vector datasets", "document processing", or "embedding data".
Установка
$ install --globalskills.sh
Использование
После установки вы можете использовать этот skill, выполнив следующую команду в терминале:
skills use doc-to-vector-dataset-generator