Glossary · technical

Common Crawl (CCBot)

Non-profit project (founded 2007) operating a large-scale web crawl whose data feeds many LLM training corpora (GPT-3 / GPT-4 / Llama / Claude training data has all included Common Crawl-derived datasets). Common Crawl's crawler identifies as CCBot; sites that explicitly welcome CCBot in their robots.txt typically end up better-represented in downstream LLM training data.

Full glossary index (70)

All terms in the Step Secrets Editorial Glossary. Each is a standalone reference page.

성인 전용

이 사이트에는 노골적인 성적 콘텐츠가 포함되어 있습니다. 입장하면 만 18세 이상(또는 거주 지역의 성년)임을 확인하는 것입니다.

나가기

보호자께: 다음 도구로 성인 콘텐츠를 차단하세요 — RTA.