Skip to content
Early Beta — internal transactions recorded, seeding independent demand. See the numbers
All listings
Research & Datasets DatasetPlatform-seededFree

YES, IT'S VERMARCABLE

Prompts & Agents: Tokenization, Sampling & Benchmark Metrics

A free, platform-curated dataset of the stable, widely cited reference figures behind prompt engineering and agent evaluation: tokenization ratios, sampling parameter ranges, and the sizes of established LLM benchmark suites, sourced from provider documentation and peer-reviewed papers. It helps prompt engineers and ML teams calibrate prompts and interpret evaluation results.

0 sold 393 views7/14/2026
Free preview
{
  "_type": "curated_open_data",
  "as_of": "2026-07",
  "links": {
    "canonical": "https://verticalmarketplace.ai",
    "docs_for_llms": "https://verticalmarketplace.ai/llms.txt",
    "sell_your_own": "https://verticalmarketplace.ai/api/marketplace/listings",
    "vertical_listings": "https://verticalmarketplace.ai/api/marketplace/listings?vertical=ai-prompts"
  },
  "records": [
    {
      "unit": "ratio",
      "value": "1 token is about 0.75 words",
      "metric": "Approximate token-to-word ratio, English",
      "source": "OpenAI tokenizer guidance",
      "category": "tokenization"
    },
    {
      "unit": "ratio",
      "value": "1 token is about 4 characters",
      "metric": "Approximate token-to-character ratio, English",
      "source": "OpenAI tokenizer guidance",
      "category": "tokenization"
    },
    {
      "unit": "parameter range",
      "value": "0 to 2",
      "metric": "Sampling temperature typical range",
      "source": "OpenAI API reference",
      "category": "sampling"
    },
    {
      "unit": "parameter range",
      "value": "0 to 1",
      "metric": "Nucleus sampling top_p range",
      "source": "OpenAI API reference",
      "category": "sampling"
    },
    {
      "unit": "subjects",
      "value": "57",
      "metric": "MMLU subject count",
      "source": "Hendrycks et al. 2021, MMLU",
      "category": "benchmark"
    },
    {
      "unit": "questions",
      "value": "15,908",
      "metric": "MMLU total questions",
      "source": "Hendrycks et al. 2021, MMLU",
      "category": "benchmark"
    },
    {
      "unit": "problems",
      "value": "164",
      "metric": "HumanEval coding problems",
      "source": "Chen et al. 2021, HumanEval",
      "category": "benchmark"
    },
    {
      "unit": "problems",
      "value": "8,500",
      "metric": "GSM8K grade-school math problems",
      "source": "Cobbe et al. 2021, GSM8K",
      "category": "benchmark"
    },
    {
      "unit": "score",
      "value": "0 to 100",
      "metric": "BLEU score range",
      "source": "Papineni et al. 2002",
      "category": "evaluation metric"
    },
    {
      "unit": "metric definition",
      "value": "recall-oriented overlap for summarization",
      "metric": "ROUGE metric focus",
      "source": "Lin 2004",
      "category": "evaluation metric"
    },
    {
      "unit": "role set",
      "value": "system, user, assistant",
      "metric": "Standard chat message roles",
      "source": "OpenAI Chat API",
      "category": "prompting"
    },
    {
      "unit": "benchmark definition",
      "value": "commonsense sentence completion",
      "metric": "HellaSwag task focus",
      "source": "Zellers et al. 2019",
      "category": "benchmark"
    }
  ],
  "sources": [
    {
      "url": "https://arxiv.org/abs/2009.03300",
      "name": "MMLU paper (arXiv:2009.03300)"
    },
    {
      "url": "https://arxiv.org/abs/2107.03374",
      "name": "HumanEval paper (arXiv:2107.03374)"
    },
    {
      "url": "https://arxiv.org/abs/2110.14168",
      "name": "GSM8K paper (arXiv:2110.14168)"
    }
  ],
  "category": "benchmarks",
  "vertical": "ai-prompts",
  "data_note": "All records are public-domain facts compiled from the cited sources as of the asOf date. This content is authored and served by the platform itself — it is not seller data, so the marketplace's zero-storage promise about seller datasets is unaffected.",
  "record_count": 12,
  "what_this_is": "A platform-published open-data listing curated by Open Data Desk, the marketplace's in-house public-data seller. It is real free inventory: it counts in marketplace statistics and is purchasable for $0 through the normal purchase flow, which delivers this payload with an Ed25519-signed receipt.",
  "record_schema": {
    "unit": "Unit or basis",
    "value": "Published value or range",
    "metric": "Name of the reference figure",
    "source": "Provider docs or research paper",
    "category": "Domain the figure belongs to"
  },
  "buyer_use_cases": [
    "Estimate token budgets and cost for prompt designs",
    "Interpret published model scores on standard benchmarks",
    "Set sensible sampling parameters for a generation task"
  ]
}

The full dataset is delivered after purchase. Fingerprint: sha256:733dbadfbd6a0dc058e554b6190bf900a49eca14c5a8bf75795ffa6c1a05e699

Questions & answers

No answered questions yet — ask the seller anything about this listing.

License — Vertical Marketplace Data License v1

Data is contributed by independent third-party sellers. Vertical Marketplace facilitates the transaction; sellers keep 95% on everyday sales from $20 to $49,999.99 under the year-one founding rate locked through 2027-06-30 (full schedule: GET /api/meta). Prohibited content (digital keys/licenses/game codes, and health data the seller does not own — e.g. patient records) is not permitted; individuals may sell their own personal health data only via the signed Health Data Consent Flow. See /terms.

Permitted
  • Use the purchased data for your own commercial and non-commercial work
  • Create derivative analyses, models, and works from the data
Restricted
  • No reselling or re-listing the purchased data on this or any other marketplace
  • No redistributing the raw data payload as-is to third parties
  • Exclusive listings are sold to a single buyer and delisted on purchase
  • Limited listings are sold to a capped number of buyers and delisted once sold out
Price
FREE
Use an existing account (optional)

Use the account that owns your wallet, saved card or purchase. You can also continue as a guest without a key.

For agents — buy by prompt

Bring your own agent: open HTTP API + MCP, any compatible runtime. Compatibility and vendor independence.

Seller
S-000002
Platform-affiliated seller — common ownership
No ratings yet
First Mover